Raster and remote sensing: geospatial agent capability
Computing over gridded data: satellite retrievals such as land surface temperature, NDVI and burn ratios, reclassifying and combining rasters, terrain and hydrology from elevation models, and values sampled from a raster. A comparison of two dates belongs here.
75 scored tasks are of this kind. All kinds of work and task formats.
Models7 of 7 models
Score
Raster and remote sensing score
Average pass@1 over 75 tasks, 0 to 100 · Higher is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
How to read this chart
What this metric means. For each of the 75 tasks whose main kind of work is raster and remote sensing, the share of a model's attempts that passed every check; then the plain average over those tasks. Tasks with no decided attempt yet are left out. These scores do not enter the Geospatial Agent Index.
Data table: Raster and remote sensing score
| Model | Creator | Raster and remote sensing score | Range while attempts are pending | Coverage (attempts) |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 66 | No pending judgements; coverage incomplete | 67 of 225 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 75 | No pending judgements; coverage incomplete | 68 of 225 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 51 | 51 to 51 | 67 of 225 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 39 | No pending judgements; coverage incomplete | 64 of 225 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 52 | No pending judgements; coverage incomplete | 73 of 225 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 74 | No pending judgements; coverage incomplete | 74 of 225 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 36 | No pending judgements; coverage incomplete | 56 of 225 planned attempts |
| Model | Score | Tasks with a decided attempt |
|---|---|---|
| Terminus-2 – DeepSeek V4 Pro (high) | 75 | 68 of 75 |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 74 | 74 of 75 |
| Terminus-2 – DeepSeek V4 Flash (high) | 66 | 67 of 75 |
| Terminus-2 – gpt-oss-120b (high) | 52 | 73 of 75 |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 51 | 63 of 75 |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 39 | 64 of 75 |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 36 | 56 of 75 |
Where the tasks come from
- GeoAgentBench: 24 tasks
- GeoBenchX: 15 tasks
- Earth-Bench: 31 tasks
- GeoAnalystBench: 5 tasks
22 other tasks also need this kind of work but count towards another kind, so they are not in this score. How each task was assigned: kinds of work.
Explore kinds of work
- Vector and overlayBuffers, spatial joins, overlay and dissolve, 18 tasks
- Networks and routingShortest paths, service areas and flows on a road network, 5 tasks
- Climate and time seriesSeries over many dates and gridded climate data, 24 tasks
- Spatial statistics and interpolationKriging, density, hot spots, regression and clustering, 24 tasks
- Mapping and cartographyMaps as the main deliverable, 65 tasks
- Recognising infeasible requestsSaying no when the data cannot answer, 79 tasks
