Source paper
the peer-reviewed or arXiv publication that defines the benchmark
Geospatial agent benchmark · not-downloaded
Also known as: GABench, GeoAgentBench, GeoAgentBench
Evaluates: complete agent systems
Tool-use execution, parameter inference, and map-product verification. GPT-4o approximately 72%, Claude 3.5 Sonnet approximately 68%; chained 4+ tool tasks drop to approximately 45%. PEA measures parameter-level accuracy within tolerance.
the peer-reviewed or arXiv publication that defines the benchmark
github.com/geox-lab/GABench%20(Git%20LFS%20req…
the published dataset the tasks are drawn from
the official repository with harness, tasks and verifiers
In shortAgents run real code inside this benchmark, and answers are checked with model judgement involved, so plausibility can pass. Dataset and code are published, so you can verify claims yourself. Production deployment, guardrails, operational correctness at scale, or data-volume awareness.
Tool-use execution, parameter inference, and map-product verification. GPT-4o approximately 72%, Claude 3.5 Sonnet approximately 68%; chained 4+ tool tasks drop to approximately 45%. PEA measures parameter-level accuracy within tolerance.
Production deployment, guardrails, operational correctness at scale, or data-volume awareness.
not-downloaded — Not downloaded (LFS assets needed). Strong comparator for execution, parameter inference, and map-product verification. Not first-ten ready until LFS assets, task CSV, GT maps, data, dependency lock, output isolation, and VLM judge posture are imported.
This entry is available as JSON at /api/registry#gabench-geoagentbench. See the registry endpoint.