Vector and overlay: geospatial agent capability
Geometry operations on points, lines and polygons: buffers, spatial selection and joins, intersection and difference, dissolve, reprojection, and counting points in polygons or grids.
18 scored tasks are of this kind. All kinds of work and task formats.
Models7 of 7 models
Score
Vector and overlay score
Average pass@1 over 18 tasks, 0 to 100 · Higher is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
How to read this chart
What this metric means. For each of the 18 tasks whose main kind of work is vector and overlay, the share of a model's attempts that passed every check; then the plain average over those tasks. Tasks with no decided attempt yet are left out. These scores do not enter the Geospatial Agent Index.
Data table: Vector and overlay score
| Model | Creator | Vector and overlay score | Range while attempts are pending | Coverage (attempts) |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 73 | No pending judgements; coverage incomplete | 15 of 54 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 71 | No pending judgements; coverage incomplete | 17 of 54 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 79 | No pending judgements; coverage incomplete | 14 of 54 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 36 | No pending judgements; coverage incomplete | 14 of 54 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 56 | No pending judgements; coverage incomplete | 18 of 54 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 78 | No pending judgements; coverage incomplete | 18 of 54 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 78 | No pending judgements; coverage incomplete | 9 of 54 planned attempts |
| Model | Score | Tasks with a decided attempt |
|---|---|---|
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 79 | 14 of 18 |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 78 | 18 of 18 |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 78 | 9 of 18 |
| Terminus-2 – DeepSeek V4 Flash (high) | 73 | 15 of 18 |
| Terminus-2 – DeepSeek V4 Pro (high) | 71 | 17 of 18 |
| Terminus-2 – gpt-oss-120b (high) | 56 | 18 of 18 |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 36 | 14 of 18 |
Where the tasks come from
- GeoAgentBench: 8 tasks
- GeoBenchX: 6 tasks
- GeoAnalystBench: 4 tasks
32 other tasks also need this kind of work but count towards another kind, so they are not in this score. How each task was assigned: kinds of work.
Explore kinds of work
- Raster and remote sensingSatellite band maths, reclassification, terrain and raster overlay, 75 tasks
- Networks and routingShortest paths, service areas and flows on a road network, 5 tasks
- Climate and time seriesSeries over many dates and gridded climate data, 24 tasks
- Spatial statistics and interpolationKriging, density, hot spots, regression and clustering, 24 tasks
- Mapping and cartographyMaps as the main deliverable, 65 tasks
- Recognising infeasible requestsSaying no when the data cannot answer, 79 tasks
