Geospatial agent benchmark · not-downloaded

GABench / GeoAgentBench

Also known as: GABench, GeoAgentBench, GeoAgentBench

Evaluates: complete agent systems

Tool-use execution, parameter inference, and map-product verification. GPT-4o approximately 72%, Claude 3.5 Sonnet approximately 68%; chained 4+ tool tasks drop to approximately 45%. PEA measures parameter-level accuracy within tolerance.

At a glance

53Tasks
YesAgents run real code
YesData + code public
NoChecked by recomputation
Data licence

Sources

every claim on this page traces to these records

In shortAgents run real code inside this benchmark, and answers are checked with model judgement involved, so plausibility can pass. Dataset and code are published, so you can verify claims yourself. Production deployment, guardrails, operational correctness at scale, or data-volume awareness.

Evidence boundaries

what a score on this benchmark does and does not tell you

A good score shows

Tool-use execution, parameter inference, and map-product verification. GPT-4o approximately 72%, Claude 3.5 Sonnet approximately 68%; chained 4+ tool tasks drop to approximately 45%. PEA measures parameter-level accuracy within tolerance.

A good score does not show

Production deployment, guardrails, operational correctness at scale, or data-volume awareness.

Arena status

not-downloaded — Not downloaded (LFS assets needed). Strong comparator for execution, parameter inference, and map-product verification. Not first-ten ready until LFS assets, task CSV, GT maps, data, dependency lock, output isolation, and VLM judge posture are imported.

All recorded facts (22 fields)
Evaluates
complete-system
Task families
buffer analysis,overlay,topology repair,raster algebra,geocoding,coordinate reprojection,map rendering
Task count
53
Difficulty
mixed (single-tool tasks through chained 4+ tool workflows)
Geography
varied (real GIS workflows across 6 GIS domains)
Modality
mixed
Input types
vector,raster,tabular,netCDF,graph or network,rendered map outputs
Output types
rendered GT maps,data products,toolchain execution traces
Execution environment
tool-augmented agent execution (117 atomic GIS tools)
Ground truth method
generated and verified GT maps and data products; ground-truth tool call sequences per task
Verifier method
TAO (Tool Accuracy Overall), TIO (Tool Invocation Overall), TEM (Tool Execution Metric), PEA (Parameter Execution Accuracy) with last-attempt alignment and file-existence checks; VLM-as-judge map comparison; execution-efficiency metrics
Deterministic verification
no — model judgement involved
Scoring dimensions
TAO,TIO,TEM,PEA,VLM map comparison,execution efficiency
Aggregation formula
composite of TAO, TIO, TEM, PEA with last-attempt alignment; file-existence checks gate pass; VLM judge for map comparison
Trials policy
pending verification (7 LLMs evaluated; GPT-4o and Claude 3.5 Sonnet perform best)
Variance reporting
pending verification
Data licence
pending verification
Code licence
pending verification
Contamination concerns
public repo uses Git LFS; normal clone without LFS leaves core task and data files as pointers, blocking reproduction
Reproducibility status
medium (repository public but LFS assets required; task CSV, GT maps, data, dependency lock, output isolation, and VLM judge posture must be imported)
First published
2026-04
Latest known update
2026-06-12

Machine-readable record

This entry is available as JSON at /api/registry#gabench-geoagentbench. See the registry endpoint.