Source paper
the peer-reviewed or arXiv publication that defines the benchmark
Geospatial agent benchmark · not-downloaded
Also known as: MapQA
Evaluates: model answers only
The benchmark reports a spatial reasoning gap: GPT-4 scores 64.2% overall and 41.3% on multi-hop (3+ steps). Tool augmentation adds 17.5 percentage points, to 81.7%. On routing, GPT-4 scores 22.4% versus 76.8% with tools.
the peer-reviewed or arXiv publication that defines the benchmark
github.com/knowledge-computing/MapQA-dataset
the published dataset the tasks are drawn from
github.com/knowledge-computing/MapQA-dataset
the official repository with harness, tasks and verifiers
In shortAgents do not run code inside this benchmark, and answers are checked by recomputation, not by another model's opinion. Dataset and code are published, so you can verify claims yourself. Code execution or artifact correctness. Does not test workflow generation, spatial data processing, or map production.
The benchmark reports a spatial reasoning gap: GPT-4 scores 64.2% overall and 41.3% on multi-hop (3+ steps). Tool augmentation adds 17.5 percentage points, to 81.7%. On routing, GPT-4 scores 22.4% versus 76.8% with tools.
Code execution or artifact correctness. Does not test workflow generation, spatial data processing, or map production.
not-downloaded — Not downloaded. Spatial reasoning and tool-augmentation comparator.
This entry is available as JSON at /api/registry#mapqa. See the registry endpoint.