Geospatial agent benchmark · downloaded

GeoAnalystBench / GISclaw

Also known as: GeoAnalystBench, GISclaw

Evaluates: end-to-end agents

Code execution success on greenfield GIS analysis. SA (single-agent ReAct) dominates DA (Plan-Execute-Replan) for strong models: DeepSeek-V3.2 96% SA vs 32% DA. Error recovery adds approximately 8 percentage points.

At a glance

50Tasks
YesAgents run real code
YesData + code public
YesChecked by recomputation
Data licence

Sources

every claim on this page traces to these records

In shortAgents run real code inside this benchmark, and answers are checked by recomputation, not by another model's opinion. Dataset and code are published, so you can verify claims yourself. Production deployment, guardrail compliance, migration nuance, documentation quality, operational correctness at scale, or data-volume awareness.

Evidence boundaries

what a score on this benchmark does and does not tell you

A good score shows

Code execution success on greenfield GIS analysis. SA (single-agent ReAct) dominates DA (Plan-Execute-Replan) for strong models: DeepSeek-V3.2 96% SA vs 32% DA. Error recovery adds approximately 8 percentage points.

A good score does not show

Production deployment, guardrail compliance, migration nuance, documentation quality, operational correctness at scale, or data-volume awareness.

Arena status

downloaded — extracted-full/ at .external/geoanalystbench/, 50 task folders with datasets + reference .py + expected outputs. GeoAnalystBench.csv. SHA-256 of full zip fdaea3ae439169174278fc22f9a3bfa0c60811c3e7f3b462c6ab4306b95e52a3.

All recorded facts (22 fields)
Evaluates
agent
Task families
vector analysis,raster processing,table operations,visualisation,network analysis,ArcGIS layer/project handling
Task count
50
Difficulty
mixed (3-10 subtasks per task, average 5.8 steps)
Geography
varied (real-world geoprocessing tasks from GIS platforms, software, online tutorials, and academic literature)
Modality
mixed
Input types
vector,raster,tabular,netCDF,network,ArcGIS layer/project files
Output types
Python code,PNG images,GeoJSON,rendered map outputs
Execution environment
persistent Python sandbox with open-source GIS libraries
Ground truth method
expert-designed reference Python solution per task; expected outputs (PNG/GeoJSON)
Verifier method
L1 API F1, L2 reasoning similarity, L3 output verification (vector geometry diff via shapely, raster pixel diff via numpy, tabular value diff via pandas). Binary pass/fail.
Deterministic verification
yes
Scoring dimensions
L1 API F1,L2 reasoning similarity,L3 output verification
Aggregation formula
binary pass/fail per task (all three layers must pass)
Trials policy
1,800 controlled runs (50 tasks x 6 backends x 2 architectures x 3 repeats)
Variance reporting
not explicitly reported
Data licence
pending verification
Code licence
pending verification (open-source repository)
Contamination concerns
tasks are expert-designed and public via CSV; contamination possible if models trained on the benchmark data
Reproducibility status
high (GeoAnalystBench is public with data and reference scripts)
First published
2026-03
Latest known update
2026-06-12
where this benchmark overlaps Arena's own task pack

Machine-readable record

This entry is available as JSON at /api/registry#geoanalystbench-gisclaw. See the registry endpoint.