Source paper
the peer-reviewed or arXiv publication that defines the benchmark
Geospatial agent benchmark · downloaded
Also known as: GeoAnalystBench, GISclaw
Evaluates: end-to-end agents
Code execution success on greenfield GIS analysis. SA (single-agent ReAct) dominates DA (Plan-Execute-Replan) for strong models: DeepSeek-V3.2 96% SA vs 32% DA. Error recovery adds approximately 8 percentage points.
the peer-reviewed or arXiv publication that defines the benchmark
github.com/GeoDS/GeoAnalystBench/blob/master/d…
the published dataset the tasks are drawn from
github.com/GeoDS/GeoAnalystBench
the official repository with harness, tasks and verifiers
In shortAgents run real code inside this benchmark, and answers are checked by recomputation, not by another model's opinion. Dataset and code are published, so you can verify claims yourself. Production deployment, guardrail compliance, migration nuance, documentation quality, operational correctness at scale, or data-volume awareness.
Code execution success on greenfield GIS analysis. SA (single-agent ReAct) dominates DA (Plan-Execute-Replan) for strong models: DeepSeek-V3.2 96% SA vs 32% DA. Error recovery adds approximately 8 percentage points.
Production deployment, guardrail compliance, migration nuance, documentation quality, operational correctness at scale, or data-volume awareness.
downloaded — extracted-full/ at .external/geoanalystbench/, 50 task folders with datasets + reference .py + expected outputs. GeoAnalystBench.csv. SHA-256 of full zip fdaea3ae439169174278fc22f9a3bfa0c60811c3e7f3b462c6ab4306b95e52a3.
This entry is available as JSON at /api/registry#geoanalystbench-gisclaw. See the registry endpoint.