Recognising infeasible requests: geospatial agent capability
GeoBenchX tasks that cannot be done with the data provided, where the only passing answer is to reject the task. Every GeoBenchX task asks for the same output files, so the instruction does not reveal which tasks these are.
79 scored tasks are of this kind. All kinds of work and task formats.
Models7 of 7 models
Score
Recognising infeasible requests score
Average pass@1 over 79 tasks, 0 to 100 · Higher is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
How to read this chart
What this metric means. For each of the 79 tasks whose main kind of work is recognising infeasible requests, the share of a model's attempts that passed every check; then the plain average over those tasks. Tasks with no decided attempt yet are left out. These scores do not enter the Geospatial Agent Index.
Data table: Recognising infeasible requests score
| Model | Creator | Recognising infeasible requests score | Range while attempts are pending | Coverage (attempts) |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 35 | 33 to 37 | 26 of 237 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 41 | No pending judgements; coverage incomplete | 39 of 237 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 44 | No pending judgements; coverage incomplete | 9 of 237 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 36 | No pending judgements; coverage incomplete | 11 of 237 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 47 | No pending judgements; coverage incomplete | 62 of 237 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 41 | No pending judgements; coverage incomplete | 63 of 237 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | No data | No data | 0 of 237 planned attempts |
| Model | Score | Tasks with a decided attempt |
|---|---|---|
| Terminus-2 – gpt-oss-120b (high) | 47 | 62 of 79 |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 44 | 9 of 79 |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 41 | 63 of 79 |
| Terminus-2 – DeepSeek V4 Pro (high) | 41 | 39 of 79 |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 36 | 11 of 79 |
| Terminus-2 – DeepSeek V4 Flash (high) | 35 | 26 of 79 |
| Terminus-2 – Qwen 3.8 27B (xhigh) | No data | 0 of 79 |
Where the tasks come from
- GeoBenchX: 79 tasks
5 other tasks also need this kind of work but count towards another kind, so they are not in this score. How each task was assigned: kinds of work.
Explore kinds of work
- Raster and remote sensingSatellite band maths, reclassification, terrain and raster overlay, 75 tasks
- Vector and overlayBuffers, spatial joins, overlay and dissolve, 18 tasks
- Networks and routingShortest paths, service areas and flows on a road network, 5 tasks
- Climate and time seriesSeries over many dates and gridded climate data, 24 tasks
- Spatial statistics and interpolationKriging, density, hot spots, regression and clustering, 24 tasks
- Mapping and cartographyMaps as the main deliverable, 65 tasks
