Axis Spatial Arena

Geospatial AI benchmarks, results and limits.

Find benchmarks relevant to your spatial work. See what they test, how they check answers and what their results can support. Read Axis Spatial’s own evaluations of specific agent configurations on defined tasks.

Selected benchmarks

source-reviewed records

Start with these reviewed benchmark records. Each explains the task, available software and data, checks, and limits of comparison.

Selected · Sources reviewed 24 Aug 2026

GeoAnalystBench / GISclaw

Executable GIS analysis with reference code and deterministic checks of vector, raster and tabular outputs.

Checks: L1 API F1, L2 reasoning similarity, L3 output verification (vector geometry diff via shapely, raster pixel diff via numpy, tabular value diff via pandas). Binary pass/fail.

Read the benchmark record

Selected · Sources reviewed 24 Aug 2026

MapQA

Spatial reasoning questions grounded in OpenStreetMap, with a separate tool-assisted mode for routing and spatial queries.

Checks: accuracy (exact match against ground-truth answer)

Read the benchmark record

Selected · Sources reviewed 24 Aug 2026

GeoNatureAgent Benchmark

Multi-turn environmental analysis through a defined tool API, checked with eight mechanistic tests per case.

Checks: eight mechanistic checks per case, including expected tools, required or forbidden text, numeric tolerance, chart generation, and loop limits

Read the benchmark record

Selected · Sources reviewed 24 Aug 2026

EO-Gym

Stateful Earth-observation work across executable tools, temporal retrieval and cross-modal evidence.

Checks: Pass@k, Tool-any, illegal-call rate, and task-category scoring

Read the benchmark record

From our evaluations

Run history · Task definitions

What happened when we ran specific agent configurations on spatial tasks. Each record explains the setup, checks and limits of the result.

Recorded evaluation · 12 Jul 2026

Vector buffer, 1 km

Across four recorded runs, the best fully accepted run passed all trials; the most recent passed three of four. A result fails if the buffer uses degrees, returns the wrong feature count or falls outside the area tolerance.

See task results

Recorded evaluation · 12 Jul 2026

Raster zonal mean

Across three recorded runs, one trial was accepted. The latest run passed one of four trials. Returning a whole-grid mean instead of the zonal mean—or inventing a number—fails the check.

See task results

Complete benchmark catalogue

reviewed records and open questions

Browse all indexed benchmark records, including records awaiting a source review. Filter by task, input and output, execution environment, and checking method.

Browse the complete catalogue

Applied Work

Axis Spatial builds spatial AI systems for organisations. Our evaluation work helps us assess the models, tools and configurations used in those systems.