Source paper
the peer-reviewed or arXiv publication that defines the benchmark
Geospatial agent benchmark · not-downloaded
Also known as: MultiGlobeQA Benchmark
Evaluates: models with tool access
Multilingual, globally stratified spatial reasoning and agentic spatial computation with execution-based checks.
the peer-reviewed or arXiv publication that defines the benchmark
huggingface.co/datasets/aiana94/MultiGlobeQA
the published dataset the tasks are drawn from
github.com/andreeaiana/MultiGlobeQA
the official repository with harness, tasks and verifiers
In shortAgents run real code inside this benchmark, and answers are checked by recomputation, not by another model's opinion. Dataset and code are published, so you can verify claims yourself. Rights to redistribute all knowledge-graph snapshots, general GIS workflow execution, or production deployment.
Multilingual, globally stratified spatial reasoning and agentic spatial computation with execution-based checks.
Rights to redistribute all knowledge-graph snapshots, general GIS workflow execution, or production deployment.
not-downloaded — Public paper, repository, and dataset reviewed; no Arena import or reproduction recorded.
This entry is available as JSON at /api/registry#multiglobeqa. See the registry endpoint.