Capabilities: geospatial agent scores by task format and kind of work
The same attempts as the Geospatial Agent Index, grouped two ways instead of by source benchmark: by task format, what the agent must deliver, and by the kind of geospatial work each task needs. Each of the 290 scored tasks has one format and one main kind of work, so a model's strengths and gaps show across benchmarks. These scores do not enter the index.
Models7 of 7 models
By task format
- Multi-step analysis (151 tasks): carry out a workflow, deliver data files and maps.
- Single-answer questions (60 tasks): one number, choice or list.
- Recognising infeasible requests (79 tasks): decline when the data cannot answer.
| Model | Multi-step analysis (151 tasks) | Single-answer questions (60 tasks) | Recognising infeasible requests (79 tasks) | Geospatial Agent Index |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Pro (high) | 84 | 45 | 41 | 73* |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 79 | 60 | 41 | 73* |
| Terminus-2 – DeepSeek V4 Flash (high) | 88 | 40 | 35 | 70* |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 78 | 32 | 44 | 60* |
| Terminus-2 – gpt-oss-120b (high) | 52 | 46 | 47 | 53* |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 41 | 47 | 36 | 43* |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 64 | 27 | No data | 62* |
Multi-step analysis score
Average pass@1 over 151 tasks, 0 to 100 · Higher is better
Data table: Multi-step analysis score
| Model | Creator | Multi-step analysis score | Range while attempts are pending | Coverage (attempts) |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 88 | 87 to 88 | 90 of 453 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 84 | No pending judgements; coverage incomplete | 101 of 453 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 78 | No pending judgements; coverage incomplete | 78 of 453 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 41 | 40 to 41 | 81 of 453 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 52 | No pending judgements; coverage incomplete | 128 of 453 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 79 | No pending judgements; coverage incomplete | 126 of 453 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 64 | No pending judgements; coverage incomplete | 53 of 453 planned attempts |
Single-answer questions score
Average pass@1 over 60 tasks, 0 to 100 · Higher is better
Data table: Single-answer questions score
| Model | Creator | Single-answer questions score | Range while attempts are pending | Coverage (attempts) |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 40 | No pending judgements; coverage incomplete | 53 of 180 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 45 | No pending judgements; coverage incomplete | 55 of 180 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 32 | 32 to 33 | 58 of 180 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 47 | No pending judgements; coverage incomplete | 51 of 180 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 46 | No pending judgements; coverage incomplete | 59 of 180 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 60 | No pending judgements; coverage incomplete | 60 of 180 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 27 | No pending judgements; coverage incomplete | 48 of 180 planned attempts |
Recognising infeasible requests score
Average pass@1 over 79 tasks, 0 to 100 · Higher is better
Data table: Recognising infeasible requests score
| Model | Creator | Recognising infeasible requests score | Range while attempts are pending | Coverage (attempts) |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 35 | 33 to 37 | 26 of 237 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 41 | No pending judgements; coverage incomplete | 39 of 237 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 44 | No pending judgements; coverage incomplete | 9 of 237 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 36 | No pending judgements; coverage incomplete | 11 of 237 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 47 | No pending judgements; coverage incomplete | 62 of 237 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 41 | No pending judgements; coverage incomplete | 63 of 237 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | No data | No data | 0 of 237 planned attempts |
By kind of work
| Model | Raster and remote sensing (75 tasks) | Vector and overlay (18 tasks) | Networks and routing (5 tasks) | Climate and time series (24 tasks) | Spatial statistics and interpolation (24 tasks) | Mapping and cartography (65 tasks) | Recognising infeasible requests (79 tasks) | Geospatial Agent Index |
|---|---|---|---|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Pro (high) | 75 | 71 | 100 | 62 | 56 | 69 | 41 | 73* |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 74 | 78 | 80 | 79 | 63 | 70 | 41 | 73* |
| Terminus-2 – DeepSeek V4 Flash (high) | 66 | 73 | 100 | 67 | 73 | 76 | 35 | 70* |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 51 | 79 | 100 | 48 | 75 | 83 | 44 | 60* |
| Terminus-2 – gpt-oss-120b (high) | 52 | 56 | 100 | 62 | 30 | 43 | 47 | 53* |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 39 | 36 | 20 | 54 | 46 | 58 | 36 | 43* |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 36 | 78 | 67 | 43 | 62 | 100 | No data | 62* |
Raster and remote sensing score
Average pass@1 over 75 tasks, 0 to 100 · Higher is better
Data table: Raster and remote sensing score
| Model | Creator | Raster and remote sensing score | Range while attempts are pending | Coverage (attempts) |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 66 | No pending judgements; coverage incomplete | 67 of 225 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 75 | No pending judgements; coverage incomplete | 68 of 225 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 51 | 51 to 51 | 67 of 225 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 39 | No pending judgements; coverage incomplete | 64 of 225 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 52 | No pending judgements; coverage incomplete | 73 of 225 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 74 | No pending judgements; coverage incomplete | 74 of 225 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 36 | No pending judgements; coverage incomplete | 56 of 225 planned attempts |
Vector and overlay score
Average pass@1 over 18 tasks, 0 to 100 · Higher is better
Data table: Vector and overlay score
| Model | Creator | Vector and overlay score | Range while attempts are pending | Coverage (attempts) |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 73 | No pending judgements; coverage incomplete | 15 of 54 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 71 | No pending judgements; coverage incomplete | 17 of 54 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 79 | No pending judgements; coverage incomplete | 14 of 54 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 36 | No pending judgements; coverage incomplete | 14 of 54 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 56 | No pending judgements; coverage incomplete | 18 of 54 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 78 | No pending judgements; coverage incomplete | 18 of 54 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 78 | No pending judgements; coverage incomplete | 9 of 54 planned attempts |
Networks and routing score
Average pass@1 over 5 tasks, 0 to 100 · Higher is better
Data table: Networks and routing score
| Model | Creator | Networks and routing score | Range while attempts are pending | Coverage (attempts) |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 100 | No pending judgements; coverage incomplete | 5 of 15 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 100 | No pending judgements; coverage incomplete | 5 of 15 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 100 | No pending judgements; coverage incomplete | 4 of 15 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 20 | No pending judgements; coverage incomplete | 5 of 15 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 100 | No pending judgements; coverage incomplete | 5 of 15 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 80 | No pending judgements; coverage incomplete | 5 of 15 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 67 | No pending judgements; coverage incomplete | 3 of 15 planned attempts |
Climate and time series score
Average pass@1 over 24 tasks, 0 to 100 · Higher is better
Data table: Climate and time series score
| Model | Creator | Climate and time series score | Range while attempts are pending | Coverage (attempts) |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 67 | No pending judgements; coverage incomplete | 24 of 72 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 62 | No pending judgements; coverage incomplete | 24 of 72 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 48 | No pending judgements; coverage incomplete | 27 of 72 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 54 | No pending judgements; coverage incomplete | 24 of 72 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 62 | No pending judgements; coverage incomplete | 24 of 72 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 79 | No pending judgements; coverage incomplete | 24 of 72 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 43 | No pending judgements; coverage incomplete | 21 of 72 planned attempts |
Spatial statistics and interpolation score
Average pass@1 over 24 tasks, 0 to 100 · Higher is better
Data table: Spatial statistics and interpolation score
| Model | Creator | Spatial statistics and interpolation score | Range while attempts are pending | Coverage (attempts) |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 73 | No pending judgements; coverage incomplete | 15 of 72 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 56 | No pending judgements; coverage incomplete | 16 of 72 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 75 | No pending judgements; coverage incomplete | 12 of 72 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 46 | No pending judgements; coverage incomplete | 13 of 72 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 30 | No pending judgements; coverage incomplete | 20 of 72 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 63 | No pending judgements; coverage incomplete | 19 of 72 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 62 | No pending judgements; coverage incomplete | 8 of 72 planned attempts |
Mapping and cartography score
Average pass@1 over 65 tasks, 0 to 100 · Higher is better
Data table: Mapping and cartography score
| Model | Creator | Mapping and cartography score | Range while attempts are pending | Coverage (attempts) |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 76 | 72 to 78 | 17 of 195 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 69 | No pending judgements; coverage incomplete | 26 of 195 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 83 | No pending judgements; coverage incomplete | 12 of 195 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 58 | 54 to 62 | 12 of 195 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 43 | No pending judgements; coverage incomplete | 47 of 195 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 70 | No pending judgements; coverage incomplete | 46 of 195 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 100 | No pending judgements; coverage incomplete | 4 of 195 planned attempts |
Recognising infeasible requests score
Average pass@1 over 79 tasks, 0 to 100 · Higher is better
Data table: Recognising infeasible requests score
| Model | Creator | Recognising infeasible requests score | Range while attempts are pending | Coverage (attempts) |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 35 | 33 to 37 | 26 of 237 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 41 | No pending judgements; coverage incomplete | 39 of 237 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 44 | No pending judgements; coverage incomplete | 9 of 237 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 36 | No pending judgements; coverage incomplete | 11 of 237 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 47 | No pending judgements; coverage incomplete | 62 of 237 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 41 | No pending judgements; coverage incomplete | 63 of 237 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | No data | No data | 0 of 237 planned attempts |
All kinds of work
- Raster and remote sensingSatellite band maths, reclassification, terrain and raster overlay, 75 tasks
- Vector and overlayBuffers, spatial joins, overlay and dissolve, 18 tasks
- Networks and routingShortest paths, service areas and flows on a road network, 5 tasks
- Climate and time seriesSeries over many dates and gridded climate data, 24 tasks
- Spatial statistics and interpolationKriging, density, hot spots, regression and clustering, 24 tasks
- Mapping and cartographyMaps as the main deliverable, 65 tasks
- Recognising infeasible requestsSaying no when the data cannot answer, 79 tasks
How tasks are assigned
The format follows from what each task must deliver; the kind of work is the one its checked result depends on most. Rules and borderline tasks.
