Comparison view

Which benchmark
proves what?

Each row links to a full reference page: what the benchmark evaluates, how it verifies answers, what it can show, and what it cannot. Claims marked pending verification could not be confirmed from public sources.

All benchmarks

19 entries · click a column head to sort

In shortTwo columns do most of the work. Executes code? — only these benchmarks can show that an agent's program actually runs. Deterministic verify? — only these check answers by recomputation rather than by another model's opinion. Open any benchmark for the full evidence-boundary record.

Sorting re-orders the same crawlable rows; nothing is hidden behind interaction.
BenchmarkTasksExecutes code?Deterministic verify?Open?Data licence
GeoAnalystBench / GISclaw50yesyesyes
GeoBenchX202yesnopartial
GEO-Bench-219noyespartialvaries by dataset
GEO-Bench-2 Leaderboardnoyespartial
GEOBench-VLMpendingnoyespartial
GeoMMBench / GeoMMAgent1053noyesyesCC BY 4.0
GABench / GeoAgentBench53yesnoyes
MapQA3154noyesyespending verification
ThinkGeopendingpending verificationyesno
Earth-Bench / Earth-Agentpendingpending verificationyesno
OpenEarth-Bench / OpenEarthAgentpendingpending verificationyesno
TerraBenchpendingpending verificationpending verificationno
RSRCC126000noyesno
UrbanSARFloods8879noyesno
GeoAI Agency Primitivespendingyesyesno
Axis Synthetic Migration Cases3yesyesnoAxis original synthetic
GeoNatureAgent Benchmark93yesyesyesunderlying indicator d…
EO-Gym9078yesyesyesfollows the licences o…
MultiGlobeQA46,060yesyesyessource knowledge-graph…