Networks and routing: geospatial agent capability
Work that needs a graph: shortest and least-time routes, service areas, and origin-destination costs and flows over a road network.
5 scored tasks are of this kind. All kinds of work and task formats.
Models7 of 7 models
Score
Networks and routing score
Average pass@1 over 5 tasks, 0 to 100 · Higher is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
How to read this chart
What this metric means. For each of the 5 tasks whose main kind of work is networks and routing, the share of a model's attempts that passed every check; then the plain average over those tasks. Tasks with no decided attempt yet are left out. These scores do not enter the Geospatial Agent Index.
Few tasks. Only 5 tasks are of this kind, so one task moves the score by a lot. Compare models here with care.
Data table: Networks and routing score
| Model | Creator | Networks and routing score | Range while attempts are pending | Coverage (attempts) |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 100 | No pending judgements; coverage incomplete | 5 of 15 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 100 | No pending judgements; coverage incomplete | 5 of 15 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 100 | No pending judgements; coverage incomplete | 4 of 15 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 20 | No pending judgements; coverage incomplete | 5 of 15 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 100 | No pending judgements; coverage incomplete | 5 of 15 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 80 | No pending judgements; coverage incomplete | 5 of 15 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 67 | No pending judgements; coverage incomplete | 3 of 15 planned attempts |
| Model | Score | Tasks with a decided attempt |
|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | 100 | 5 of 5 |
| Terminus-2 – DeepSeek V4 Pro (high) | 100 | 5 of 5 |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 100 | 4 of 5 |
| Terminus-2 – gpt-oss-120b (high) | 100 | 5 of 5 |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 80 | 5 of 5 |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 67 | 3 of 5 |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 20 | 5 of 5 |
Where the tasks come from
- GeoAgentBench: 5 tasks
2 other tasks also need this kind of work but count towards another kind, so they are not in this score. How each task was assigned: kinds of work.
Explore kinds of work
- Raster and remote sensingSatellite band maths, reclassification, terrain and raster overlay, 75 tasks
- Vector and overlayBuffers, spatial joins, overlay and dissolve, 18 tasks
- Climate and time seriesSeries over many dates and gridded climate data, 24 tasks
- Spatial statistics and interpolationKriging, density, hot spots, regression and clustering, 24 tasks
- Mapping and cartographyMaps as the main deliverable, 65 tasks
- Recognising infeasible requestsSaying no when the data cannot answer, 79 tasks
