Mapping and cartography: geospatial agent capability
Tasks where the map is the main deliverable and the analysis behind it is light: joining tables to boundaries, filtering, then a choropleth, bivariate or point map. The map is graded by an AI judge against a private checklist.
65 scored tasks are of this kind. All kinds of work and task formats.
Models7 of 7 models
Score
Mapping and cartography score
Average pass@1 over 65 tasks, 0 to 100 · Higher is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
How to read this chart
What this metric means. For each of the 65 tasks whose main kind of work is mapping and cartography, the share of a model's attempts that passed every check; then the plain average over those tasks. Tasks with no decided attempt yet are left out. These scores do not enter the Geospatial Agent Index.
Data table: Mapping and cartography score
| Model | Creator | Mapping and cartography score | Range while attempts are pending | Coverage (attempts) |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 76 | 72 to 78 | 17 of 195 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 69 | No pending judgements; coverage incomplete | 26 of 195 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 83 | No pending judgements; coverage incomplete | 12 of 195 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 58 | 54 to 62 | 12 of 195 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 43 | No pending judgements; coverage incomplete | 47 of 195 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 70 | No pending judgements; coverage incomplete | 46 of 195 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 100 | No pending judgements; coverage incomplete | 4 of 195 planned attempts |
| Model | Score | Tasks with a decided attempt |
|---|---|---|
| Terminus-2 – Qwen 3.8 27B (xhigh) | 100 | 4 of 65 |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 83 | 12 of 65 |
| Terminus-2 – DeepSeek V4 Flash (high) | 76 | 17 of 65 |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 70 | 46 of 65 |
| Terminus-2 – DeepSeek V4 Pro (high) | 69 | 26 of 65 |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 58 | 12 of 65 |
| Terminus-2 – gpt-oss-120b (high) | 43 | 47 of 65 |
Where the tasks come from
- GeoAgentBench: 3 tasks
- GeoBenchX: 59 tasks
- GeoAnalystBench: 3 tasks
39 other tasks also need this kind of work but count towards another kind, so they are not in this score. How each task was assigned: kinds of work.
Explore kinds of work
- Raster and remote sensingSatellite band maths, reclassification, terrain and raster overlay, 75 tasks
- Vector and overlayBuffers, spatial joins, overlay and dissolve, 18 tasks
- Networks and routingShortest paths, service areas and flows on a road network, 5 tasks
- Climate and time seriesSeries over many dates and gridded climate data, 24 tasks
- Spatial statistics and interpolationKriging, density, hot spots, regression and clustering, 24 tasks
- Recognising infeasible requestsSaying no when the data cannot answer, 79 tasks
