Spatial statistics and interpolation: geospatial agent capability
Statistical models of space: kriging and other interpolation, kernel density and heatmaps, local Moran's I hot spots, regression including geographically weighted regression, point-pattern functions and clustering.
24 scored tasks are of this kind. All kinds of work and task formats.
Models7 of 7 models
Score
Spatial statistics and interpolation score
Average pass@1 over 24 tasks, 0 to 100 · Higher is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
How to read this chart
What this metric means. For each of the 24 tasks whose main kind of work is spatial statistics and interpolation, the share of a model's attempts that passed every check; then the plain average over those tasks. Tasks with no decided attempt yet are left out. These scores do not enter the Geospatial Agent Index.
Data table: Spatial statistics and interpolation score
| Model | Creator | Spatial statistics and interpolation score | Range while attempts are pending | Coverage (attempts) |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 73 | No pending judgements; coverage incomplete | 15 of 72 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 56 | No pending judgements; coverage incomplete | 16 of 72 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 75 | No pending judgements; coverage incomplete | 12 of 72 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 46 | No pending judgements; coverage incomplete | 13 of 72 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 30 | No pending judgements; coverage incomplete | 20 of 72 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 63 | No pending judgements; coverage incomplete | 19 of 72 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 62 | No pending judgements; coverage incomplete | 8 of 72 planned attempts |
| Model | Score | Tasks with a decided attempt |
|---|---|---|
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 75 | 12 of 24 |
| Terminus-2 – DeepSeek V4 Flash (high) | 73 | 15 of 24 |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 63 | 19 of 24 |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 62 | 8 of 24 |
| Terminus-2 – DeepSeek V4 Pro (high) | 56 | 16 of 24 |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 46 | 13 of 24 |
| Terminus-2 – gpt-oss-120b (high) | 30 | 20 of 24 |
Where the tasks come from
- GeoAgentBench: 7 tasks
- GeoBenchX: 14 tasks
- GeoAnalystBench: 3 tasks
6 other tasks also need this kind of work but count towards another kind, so they are not in this score. How each task was assigned: kinds of work.
Explore kinds of work
- Raster and remote sensingSatellite band maths, reclassification, terrain and raster overlay, 75 tasks
- Vector and overlayBuffers, spatial joins, overlay and dissolve, 18 tasks
- Networks and routingShortest paths, service areas and flows on a road network, 5 tasks
- Climate and time seriesSeries over many dates and gridded climate data, 24 tasks
- Mapping and cartographyMaps as the main deliverable, 65 tasks
- Recognising infeasible requestsSaying no when the data cannot answer, 79 tasks
