Geospatial agent benchmark · downloaded

GeoBenchX

Also known as: GeoBenchX

Evaluates: models with tool access

Geospatial tool-use and agent task completion with reference-solution comparison.

At a glance

202Tasks
YesAgents run real code
PartlyData + code public
NoChecked by recomputation
Data licence

Sources

every claim on this page traces to these records

In shortAgents run real code inside this benchmark, and answers are checked with model judgement involved, so plausibility can pass. Part of the evidence base (dataset or code, not both) is published. Reference solutions are tool traces, not final numeric or geospatial gold outputs. Does not prove production deployment, guardrail compliance, or deterministic output verification.

Evidence boundaries

what a score on this benchmark does and does not tell you

A good score shows

Geospatial tool-use and agent task completion with reference-solution comparison.

A good score does not show

Reference solutions are tool traces, not final numeric or geospatial gold outputs. Does not prove production deployment, guardrail compliance, or deterministic output verification.

Arena status

downloaded — /root/axis-spatial/.external/geobenchx/ — upstream repo + geobenchx-data-download/ (159 files: GeoData, StatData, source notes, BibTeX). SHA-256 afb80eecdb731425743efaf9c916a59f54db17567ca5576e83989223cb5595e9.

All recorded facts (22 fields)
Evaluates
model+tools
Task families
geospatial tool-use,mixed reasoning,geospatial API calls
Task count
202
Difficulty
mixed
Geography
varied
Modality
mixed
Input types
tool calls,geospatial APIs,mixed reasoning prompts
Output types
tool traces,reference solutions
Execution environment
LangGraph ReAct agent with a 25-iteration limit
Ground truth method
reference solutions (tool traces, not final numeric or geospatial gold outputs)
Verifier method
LLM-as-judge plus reference-solution comparison
Deterministic verification
no — model judgement involved
Scoring dimensions
LLM judge score,reference-solution comparison
Aggregation formula
pending verification
Trials policy
pending verification
Variance reporting
pending verification
Data licence
pending verification
Code licence
MIT
Contamination concerns
reference solutions are tool traces, not final numeric or geospatial gold outputs; LLM-judge scoring is non-deterministic
Reproducibility status
medium (repository public, data via Google Drive bundle)
First published
2025-03
Latest known update
2025-11-18

Machine-readable record

This entry is available as JSON at /api/registry#geobenchx. See the registry endpoint.