Geospatial agent benchmark · not-downloaded

MapQA

Also known as: MapQA

Evaluates: model answers only

The benchmark reports a spatial reasoning gap: GPT-4 scores 64.2% overall and 41.3% on multi-hop (3+ steps). Tool augmentation adds 17.5 percentage points, to 81.7%. On routing, GPT-4 scores 22.4% versus 76.8% with tools.

At a glance

3154Tasks
NoAgents run real code
YesData + code public
YesChecked by recomputation
pending verificationData licence

Sources

every claim on this page traces to these records

In shortAgents do not run code inside this benchmark, and answers are checked by recomputation, not by another model's opinion. Dataset and code are published, so you can verify claims yourself. Code execution or artifact correctness. Does not test workflow generation, spatial data processing, or map production.

Evidence boundaries

what a score on this benchmark does and does not tell you

A good score shows

The benchmark reports a spatial reasoning gap: GPT-4 scores 64.2% overall and 41.3% on multi-hop (3+ steps). Tool augmentation adds 17.5 percentage points, to 81.7%. On routing, GPT-4 scores 22.4% versus 76.8% with tools.

A good score does not show

Code execution or artifact correctness. Does not test workflow generation, spatial data processing, or map production.

Arena status

not-downloaded — Not downloaded. Spatial reasoning and tool-augmentation comparator.

All recorded facts (22 fields)
Evaluates
model
Task families
proximity,containment,routing,aggregation,comparison
Task count
3154
Difficulty
mixed (single-hop through multi-hop 3+ step reasoning)
Geography
Southern California and Illinois
Modality
vector
Input types
OSM geometries,POI metadata,routing queries
Output types
answer (text)
Execution environment
LLM inference; optional tool-augmented mode with geocoding API, routing API, and spatial query tool
Ground truth method
derived from OSM geometries and POI metadata with known spatial relationships
Verifier method
accuracy (exact match against ground-truth answer)
Deterministic verification
yes
Scoring dimensions
accuracy
Aggregation formula
accuracy (percentage correct); per-category breakdown (proximity, containment, routing, aggregation, comparison)
Trials policy
pending verification (GPT-4, GPT-3.5, Gemini Pro, LLaMA-3-70B tested; plus tool-augmented GPT-4)
Variance reporting
pending verification
Data licence
pending verification (underlying OSM data is ODbL)
Code licence
pending verification
Contamination concerns
questions derived from public OSM data; contamination possible if models trained on similar QA pairs
Reproducibility status
medium (methodology public; dataset availability pending verification)
First published
2025-03
Latest known update
pending verification

Machine-readable record

This entry is available as JSON at /api/registry#mapqa. See the registry endpoint.