Axis Spatial · public evidence index

Geospatial AI
Benchmark Registry

The registry records what each benchmark measures and how results are checked. It also states where the evidence stops. 19 benchmarks are indexed from linked public sources. Unknown fields remain explicit.

19public benchmarks indexed
7own benchmark tasks
11runs recorded
360trials executed

Registry matrix

Registry last reviewed 24 Aug 2026 · horizontal scroll on small screens

In shortEach row is one public benchmark. Exec = agents run real code; Det. verify = answers are checked by recomputation, not by another model. Open? = dataset and code are published so you can verify claims yourself. Follow a benchmark's name for its full record: papers, data, limits, and how to read a score.

Five columns for scanning; every claim keeps its explicit unknown. Each benchmark name opens its full record page.
BenchmarkTasksExec?Det. verify?Open?
GeoAnalystBench / GISclaw50yesyesyes
GeoBenchX202yesnopartial
GEO-Bench-219noyespartial
GEO-Bench-2 Leaderboardnoyespartial
GEOBench-VLMpendingnoyespartial
GeoMMBench / GeoMMAgent1053noyesyes
GABench / GeoAgentBench53yesnoyes
MapQA3154noyesyes
ThinkGeopendingpending verificationyesno
Earth-Bench / Earth-Agentpendingpending verificationyesno
OpenEarth-Bench / OpenEarthAgentpendingpending verificationyesno
TerraBenchpendingpending verificationpending verificationno
RSRCC126000noyesno
UrbanSARFloods8879noyesno
GeoAI Agency Primitivespendingyesyesno
Axis Synthetic Migration Cases3yesyesno
GeoNatureAgent Benchmark93yesyesyes
EO-Gym9078yesyesyes
MultiGlobeQA46,060yesyesyes

How to read the registry

Use the JSON endpoint

Comparison unit

A benchmark result is meaningful only with its task set, dataset, tool surface, runtime and verifier. Model names alone are not comparable.

Verification standard

Deterministic checks and inspectable artifacts offer stronger evidence than self-reported scores. Where a benchmark relies on model judgement we say so.

Honest unknowns

Fields marked pending verification could not be confirmed from public sources. They stay visible rather than being guessed.

Arena's own results

The Arena executes its own sealed task pack and publishes every run, including blocked, inconclusive, and rejected outcomes. See the methodology.