Comparison view

Which benchmark
proves what?

Each row links to a full reference page: what the benchmark evaluates, how it verifies answers, what it can show, and what it cannot. Claims marked pending verification could not be confirmed from public sources.

All benchmarks

19 entries · click a column head to sort

In shortTwo columns do most of the work. Executes code? — only these benchmarks can show that an agent's program actually runs. Deterministic verify? — only these check answers by recomputation rather than by another model's opinion. Open any benchmark for the full evidence-boundary record.

Sorting re-orders the same crawlable rows; nothing is hidden behind interaction.
BenchmarkTasksExecutes code?Deterministic verify?Open?Data licence
GeoAnalystBench / GISclaw50yesyesyes—
GeoBenchX202yesnopartial—
GEO-Bench-219noyespartialvaries by dataset
GEO-Bench-2 Leaderboard—noyespartial—
GEOBench-VLMpendingnoyespartial—
GeoMMBench / GeoMMAgent1053noyesyesCC BY 4.0
GABench / GeoAgentBench53yesnoyes—
MapQA3154noyesyespending verification
ThinkGeopendingpending verificationyesno—
Earth-Bench / Earth-Agentpendingpending verificationyesno—
OpenEarth-Bench / OpenEarthAgentpendingpending verificationyesno—
TerraBenchpendingpending verificationpending verificationno—
RSRCC126000noyesno—
UrbanSARFloods8879noyesno—
GeoAI Agency Primitivespendingyesyesno—
Axis Synthetic Migration Cases3yesyesnoAxis original synthetic
GeoNatureAgent Benchmark93yesyesyesunderlying indicator d…
EO-Gym9078yesyesyesfollows the licences o…
MultiGlobeQA46,060yesyesyessource knowledge-graph…