Mistral AI models: geospatial agent results
Geospatial Agent Benchmarks tests 1 model configuration from Mistral AI.
Highlights
- Highest Geospatial Agent Index (ranked entries): No data yet.
- Lowest cost per task: Terminus-2 – Mistral Small 3.1 24B (Non-reasoning) ($0.247).
- Fastest per task: Terminus-2 – Mistral Small 3.1 24B (Non-reasoning) (15.6 min).
Results
Geospatial Agent Index
Equal-weight average of the evaluation scores, 0 to 100 · Higher is better
- Partial coverage
Data table: Geospatial Agent Index
| Model | Creator | Geospatial Agent Index | Range while attempts are pending | Benchmarks covered | Coverage (attempts) |
|---|---|---|---|---|---|
| Claude Code – Claude Opus 5.5 (xhigh) | Anthropic | 82 (attempts per task: 1) | No pending judgements | 4 of 4 | 290 of 290 planned attempts |
| Claude Code – Claude Opus 5.5 (high) | Anthropic | 82* (attempts per task: 2.5) | 63 to 84 | 4 of 4 | 531 of 870 planned attempts |
| Claude Code – Claude Sonnet 5.5 (high) | Anthropic | 81* (attempts per task: 2) | 77 to 81 | 4 of 4 | 556 of 870 planned attempts |
| Claude Code – Claude Opus 5.5 (medium) | Anthropic | 80* (attempts per task: 2.7) | 58 to 83 | 4 of 4 | 573 of 870 planned attempts |
| Codex – GPT-6.1 Sol (high) | OpenAI | 80 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Codex – GPT-6.1 Sol (medium) | OpenAI | 79* (attempts per task: 2.6) | No pending judgements; coverage incomplete | 4 of 4 | 749 of 870 planned attempts (6 excluded) |
| Codex – GPT-6.1 Sol (xhigh) | OpenAI | 78* (attempts per task: 2.6) | 69 to 78 | 4 of 4 | 668 of 870 planned attempts |
| Claude Code – Claude Haiku 5.5 (xhigh) | Anthropic | 78 (attempts per task: 1) | No pending judgements | 4 of 4 | 290 of 290 planned attempts |
| Codex – GPT-6 Luna (xhigh) | OpenAI | 77* (attempts per task: 2.4) | 69 to 78 | 4 of 4 | 627 of 870 planned attempts |
| Claude Code – Claude Haiku 5.5 (medium) | Anthropic | 77* (attempts per task: 2) | 74 to 78 | 4 of 4 | 538 of 870 planned attempts |
| Claude Code – Claude Haiku 5.5 (high) | Anthropic | 77* (attempts per task: 2.3) | 61 to 79 | 4 of 4 | 526 of 870 planned attempts |
| Claude Code – Claude Sonnet 5.5 (medium) | Anthropic | 76* (attempts per task: 2.2) | 71 to 76 | 4 of 4 | 580 of 870 planned attempts |
| Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high) | Anthropic | 75* (attempts per task: 1.2) | 74 to 75 | 4 of 4 | 335 of 870 planned attempts (138 excluded) |
| Terminus-2 – Gemini 3.1 Pro (preview) (high) | 75 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts | |
| Terminus-2 – DeepSeek V4 Flash (max) | DeepSeek | 74* (attempts per task: 1.2) | 73 to 75 | 4 of 4 | 344 of 870 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 74 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Claude Code – Claude Sonnet 5.5 (xhigh) | Anthropic | 74 (attempts per task: 1) | No pending judgements | 4 of 4 | 290 of 290 planned attempts |
| Terminus-2 – GLM-5.2 (max) | Z.ai | 74* (attempts per task: 1.2) | 74 to 74 | 4 of 4 | 338 of 870 planned attempts (1 excluded) |
| Terminus-2 – Kimi K2.7 Code (always on) | Moonshot AI | 74 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Terminus-2 – GLM-5.2 (high) | Z.ai | 72* (attempts per task: 1.1) | 71 to 72 | 4 of 4 | 310 of 870 planned attempts (1 excluded) |
| Terminus-2 – Kimi K2.6 (always on) | Moonshot AI | 72 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Terminus-2 – GLM-5.3 (low) | Z.ai | 72* (attempts per task: 1) | 70 to 72 | 4 of 4 | 285 of 290 planned attempts |
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 71 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Terminus-2 – GLM-5.2 (none) | Z.ai | 70* (attempts per task: 1.6) | 68 to 70 | 4 of 4 | 446 of 870 planned attempts |
| Terminus-2 – GLM-5.3 (max, default) | Z.ai | 68* (attempts per task: 1.2) | 67 to 68 | 4 of 4 | 335 of 870 planned attempts |
| Terminus-2 – GLM-5.3-Flash (low) | Z.ai | 66* (attempts per task: 1.4) | 61 to 67 | 4 of 4 | 380 of 870 planned attempts (4 excluded) |
| Terminus-2 – Gemini 3.1 Pro (preview) (low) | 65 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts (22 excluded) | |
| Terminus-2 – Gemma 4 26B A4B (thinking on) | 62 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts | |
| Terminus-2 – Gemini 3.1 Flash-Lite (high) | 59* (attempts per task: 3) | 59 to 59 | 4 of 4 | 864 of 870 planned attempts (6 excluded) | |
| Terminus-2 – GLM-5.3-Flash (max) | Z.ai | 56* (attempts per task: 1.2) | 55 to 56 | 4 of 4 | 342 of 870 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 52 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Terminus-2 – Gemini 3.1 Flash-Lite (low) | 49 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (thinking on) | Z.ai | 40 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Terminus-2 – gpt-oss-120b (low) | OpenAI | 37 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Terminus-2 – Gemini 3.1 Flash-Lite (minimal) | 34* (attempts per task: 3) | No pending judgements; coverage incomplete | 4 of 4 | 869 of 870 planned attempts | |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 62* (attempts per task: 1 on 101 of 290 tasks so far) | No pending judgements; coverage incomplete | 3 of 4 | 101 of 870 planned attempts |
| Terminus-2 – Mistral Small 3.1 24B (Non-reasoning) | Mistral AI | 0* (attempts per task: 1 on 9 of 290 tasks so far) | No pending judgements; coverage incomplete | 1 of 4 | 9 of 351 planned attempts |
| Model | Geospatial Agent Index | GeoAgentBench | GeoBenchX | Earth-Bench | GeoAnalystBench | Cost per task | Time per task | Context window |
|---|---|---|---|---|---|---|---|---|
| Terminus-2 – Mistral Small 3.1 24B (Non-reasoning) | 0* (attempts per task: 1 on 9 of 290 tasks so far) | No data | No data | 0 | No data | $0.247 | 15.6 min | 128,000 |
