Source paper
the peer-reviewed or arXiv publication that defines the benchmark
Geospatial agent benchmark · not-downloaded
Also known as: GEO-Bench-2
Evaluates: model answers only
EO foundation-model downstream task quality across 19 datasets spanning multiple sensors.
the peer-reviewed or arXiv publication that defines the benchmark
github.com/The-AI-Alliance/GEO-Bench-2
the official repository with harness, tasks and verifiers
huggingface.co/spaces/aialliance/GEO-Bench-2-L…
the maintainers' published results, where available
In shortAgents do not run code inside this benchmark, and answers are checked by recomputation, not by another model's opinion. Part of the evidence base (dataset or code, not both) is published. Agent or tool-use execution. Does not test code generation, workflow planning, or artifact production.
EO foundation-model downstream task quality across 19 datasets spanning multiple sensors.
Agent or tool-use execution. Does not test code generation, workflow planning, or artifact production.
not-downloaded — Not downloaded. Planned as model-selection benchmark and EO or GFM comparator. Pin HF dataset versions, licences, selected split, and cache refs before use.
This entry is available as JSON at /api/registry#geo-bench-2. See the registry endpoint.