Geospatial agent benchmark · not-downloaded

MultiGlobeQA

Also known as: MultiGlobeQA Benchmark

Evaluates: models with tool access

Multilingual, globally stratified spatial reasoning and agentic spatial computation with execution-based checks.

At a glance

46,060 questions per languageTasks
YesAgents run real code
YesData + code public
YesChecked by recomputation
source knowledge-graph…Data licence

Sources

every claim on this page traces to these records

In shortAgents run real code inside this benchmark, and answers are checked by recomputation, not by another model's opinion. Dataset and code are published, so you can verify claims yourself. Rights to redistribute all knowledge-graph snapshots, general GIS workflow execution, or production deployment.

Evidence boundaries

what a score on this benchmark does and does not tell you

A good score shows

Multilingual, globally stratified spatial reasoning and agentic spatial computation with execution-based checks.

A good score does not show

Rights to redistribute all knowledge-graph snapshots, general GIS workflow execution, or production deployment.

Arena status

not-downloaded — Public paper, repository, and dataset reviewed; no Arena import or reproduction recorded.

All recorded facts (22 fields)
Evaluates
model+tools
Task families
spatial functions, retrieval, spatial computation, and a multimodal slice
Task count
46,060 questions per language
Difficulty
varied; Tier-3 uses retrieval tools and a restricted Python executor
Geography
201 countries and territories
Modality
text with a multimodal slice
Input types
multilingual spatial questions, retrieval tools, and spatial SQL data
Output types
15 answer formats and executable spatial queries
Execution environment
Tier-3 agentic evaluation with retrieval tools and a restricted Python executor
Ground truth method
execution-based spatial SQL ground truth over three knowledge graphs
Verifier method
exact match, normalized error, coverage, and false-refusal scoring
Deterministic verification
yes
Scoring dimensions
exact match, normalized error, coverage, and false refusal
Aggregation formula
reported by language, spatial-function family, answer format, and evaluation tier
Trials policy
pending verification
Variance reporting
pending verification
Data licence
source knowledge-graph licences vary
Code licence
MIT (repository)
Contamination concerns
public knowledge graphs and question data; source snapshots require licence review
Reproducibility status
high for code and dataset; source knowledge-graph snapshots require review
First published
2026-08
Latest known update
2026-08-04

Machine-readable record

This entry is available as JSON at /api/registry#multiglobeqa. See the registry endpoint.