Comparison unit
A benchmark result is meaningful only with its task set, dataset, tool surface, runtime and verifier. Model names alone are not comparable.
Axis Spatial · public evidence index
The registry records what each benchmark measures and how results are checked. It also states where the evidence stops. 19 benchmarks are indexed from linked public sources. Unknown fields remain explicit.
In shortEach row is one public benchmark. Exec = agents run real code; Det. verify = answers are checked by recomputation, not by another model. Open? = dataset and code are published so you can verify claims yourself. Follow a benchmark's name for its full record: papers, data, limits, and how to read a score.
| Benchmark | Tasks | Exec? | Det. verify? | Open? |
|---|---|---|---|---|
| GeoAnalystBench / GISclaw | 50 | yes | yes | yes |
| GeoBenchX | 202 | yes | no | partial |
| GEO-Bench-2 | 19 | no | yes | partial |
| GEO-Bench-2 Leaderboard | — | no | yes | partial |
| GEOBench-VLM | pending | no | yes | partial |
| GeoMMBench / GeoMMAgent | 1053 | no | yes | yes |
| GABench / GeoAgentBench | 53 | yes | no | yes |
| MapQA | 3154 | no | yes | yes |
| ThinkGeo | pending | pending verification | yes | no |
| Earth-Bench / Earth-Agent | pending | pending verification | yes | no |
| OpenEarth-Bench / OpenEarthAgent | pending | pending verification | yes | no |
| TerraBench | pending | pending verification | pending verification | no |
| RSRCC | 126000 | no | yes | no |
| UrbanSARFloods | 8879 | no | yes | no |
| GeoAI Agency Primitives | pending | yes | yes | no |
| Axis Synthetic Migration Cases | 3 | yes | yes | no |
| GeoNatureAgent Benchmark | 93 | yes | yes | yes |
| EO-Gym | 9078 | yes | yes | yes |
| MultiGlobeQA | 46,060 | yes | yes | yes |
A benchmark result is meaningful only with its task set, dataset, tool surface, runtime and verifier. Model names alone are not comparable.
Deterministic checks and inspectable artifacts offer stronger evidence than self-reported scores. Where a benchmark relies on model judgement we say so.
Fields marked pending verification could not be confirmed from public sources. They stay visible rather than being guessed.
The Arena executes its own sealed task pack and publishes every run, including blocked, inconclusive, and rejected outcomes. See the methodology.