Leaderboard: AI models as geospatial agents
Every model configuration ranked by the Geospatial Agent Index, with its score on each benchmark, cost and time per task, and its specification. See the overview for charts.
Ongoing evaluation: results update as runs complete. Last updated 9 October 2026.
Highlights
- Geospatial Agent Index: highest is Terminus-2 – DeepSeek V4 Pro (high) (73*).
- Cost per task: lowest is Terminus-2 – Gemma 4 26B A4B (Reasoning) ($0.012).
- Time per task: fastest is Terminus-2 – gpt-oss-120b (high) (4.6 min).
- Context window: largest is Terminus-2 – DeepSeek V4 Flash (high) (1,048,576 tokens).
All models
Overall
| Rank | Model | Creator | Geospatial Agent Index | GeoAgentBench | GeoBenchX | Earth-Bench | GeoAnalystBench | Cost per task | Time per task | Tokens per attempt | Attempts decided | Context window | Price per 1M input / output tokens | Image input | Further analysis |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 73* | 98 | 45 | 48 | 100 | $0.261 | 8.6 min | 197k | 195 | 1,048,576 | $1.32 / $3.96 | No | Model |
| 2 | Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 73* | 90 | 54 | 62 | 84 | $0.126 | 6.7 min | 309k | 249 | 262,144 | $0.95 / $4.00 | Yes | Model |
| 3 | Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 70* | 96 | 44 | 40 | 100 | $0.060 | 10.0 min | 300k | 169 | 1,048,576 | $0.44 / $1.32 | No | Model |
| 4 | Terminus-2 – Gemma 4 26B A4B (Reasoning) | 60* | 86 | 55 | 30 | 68 | $0.012 | 8.6 min | 89k | 145 | 256,000 | $0.10 / $0.30 | Yes | Model | |
| 5 | Terminus-2 – gpt-oss-120b (high) | OpenAI | 53* | 70 | 42 | 46 | 53 | $0.026 | 4.6 min | 59k | 249 | 128,000 | $0.35 / $0.75 | No | Model |
| 6 | Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 43* | 36 | 46 | 48 | 42 | $0.025 | 9.5 min | 332k | 143 | 131,072 | $0.06 / $0.40 | No | Model |
| Rank | Model | Creator | Geospatial Agent Index | GeoAgentBench | GeoBenchX | Earth-Bench | GeoAnalystBench | Cost per task | Time per task | Tokens per attempt | Attempts decided | Context window | Price per 1M input / output tokens | Image input | Further analysis |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| – | Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 62* | 60 | No data | 27 | 100 | $0.066 | 19.8 min | 90k | 101 | 262,144 | $0.45 / $3.20 | Yes | Model |
By task format
| Model | Multi-step analysis (151 tasks) | Single-answer questions (60 tasks) | Recognising infeasible requests (79 tasks) | Geospatial Agent Index |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Pro (high) | 84 | 45 | 41 | 73* |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 79 | 60 | 41 | 73* |
| Terminus-2 – DeepSeek V4 Flash (high) | 88 | 40 | 35 | 70* |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 78 | 32 | 44 | 60* |
| Terminus-2 – gpt-oss-120b (high) | 52 | 46 | 47 | 53* |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 41 | 47 | 36 | 43* |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 64 | 27 | No data | 62* |
By kind of work
| Model | Raster and remote sensing (75 tasks) | Vector and overlay (18 tasks) | Networks and routing (5 tasks) | Climate and time series (24 tasks) | Spatial statistics and interpolation (24 tasks) | Mapping and cartography (65 tasks) | Recognising infeasible requests (79 tasks) | Geospatial Agent Index |
|---|---|---|---|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Pro (high) | 75 | 71 | 100 | 62 | 56 | 69 | 41 | 73* |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 74 | 78 | 80 | 79 | 63 | 70 | 41 | 73* |
| Terminus-2 – DeepSeek V4 Flash (high) | 66 | 73 | 100 | 67 | 73 | 76 | 35 | 70* |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 51 | 79 | 100 | 48 | 75 | 83 | 44 | 60* |
| Terminus-2 – gpt-oss-120b (high) | 52 | 56 | 100 | 62 | 30 | 43 | 47 | 53* |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 39 | 36 | 20 | 54 | 46 | 58 | 36 | 43* |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 36 | 78 | 67 | 43 | 62 | 100 | No data | 62* |
