Axis Spatial

Mistral AI models: geospatial agent results

Geospatial Agent Benchmarks tests 1 model configuration from Mistral AI.

Highlights

Results

Geospatial Agent Index

Equal-weight average of the evaluation scores, 0 to 100 · Higher is better

  • Partial coverage
Geospatial Agent IndexClaude Code – Claude Opus 5.5 (xhigh): 82; Claude Code – Claude Opus 5.5 (high): 82* (partial coverage); Claude Code – Claude Sonnet 5.5 (high): 81* (partial coverage); Claude Code – Claude Opus 5.5 (medium): 80* (partial coverage); Codex – GPT-6.1 Sol (high): 80; Codex – GPT-6.1 Sol (medium): 79* (partial coverage); Codex – GPT-6.1 Sol (xhigh): 78* (partial coverage); Claude Code – Claude Haiku 5.5 (xhigh): 78; Codex – GPT-6 Luna (xhigh): 77* (partial coverage); Claude Code – Claude Haiku 5.5 (medium): 77* (partial coverage); Claude Code – Claude Haiku 5.5 (high): 77* (partial coverage); Claude Code – Claude Sonnet 5.5 (medium): 76* (partial coverage); Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high): 75* (partial coverage); Terminus-2 – Gemini 3.1 Pro (preview) (high): 75; Terminus-2 – DeepSeek V4 Flash (max): 74* (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 74; Claude Code – Claude Sonnet 5.5 (xhigh): 74; Terminus-2 – GLM-5.2 (max): 74* (partial coverage); Terminus-2 – Kimi K2.7 Code (always on): 74; Terminus-2 – GLM-5.2 (high): 72* (partial coverage); Terminus-2 – Kimi K2.6 (always on): 72; Terminus-2 – GLM-5.3 (low): 72* (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 71; Terminus-2 – GLM-5.2 (none): 70* (partial coverage); Terminus-2 – GLM-5.3 (max, default): 68* (partial coverage); Terminus-2 – GLM-5.3-Flash (low): 66* (partial coverage); Terminus-2 – Gemini 3.1 Pro (preview) (low): 65; Terminus-2 – Gemma 4 26B A4B (thinking on): 62; Terminus-2 – Gemini 3.1 Flash-Lite (high): 59* (partial coverage); Terminus-2 – GLM-5.3-Flash (max): 56* (partial coverage); Terminus-2 – gpt-oss-120b (high): 52; Terminus-2 – Gemini 3.1 Flash-Lite (low): 49; Terminus-2 – GLM-4.7-Flash (thinking on): 40; Terminus-2 – gpt-oss-120b (low): 37; Terminus-2 – Gemini 3.1 Flash-Lite (minimal): 34* (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 62* (partial coverage); Terminus-2 – Mistral Small 3.1 24B (Non-reasoning): 0* (partial coverage)025507510082Claude CodeClaude Opus5.5 (xhigh)82*Claude CodeClaude Opus5.5 (high)81*Claude CodeClaude Sonnet5.5 (high)80*Claude CodeClaude Opus5.5 (medium)80CodexGPT-6.1 Sol(high)79*CodexGPT-6.1 Sol(medium)78*CodexGPT-6.1 Sol(xhigh)78Claude CodeClaude Haiku5.5 (xhigh)77*CodexGPT-6 Luna(xhigh)77*Claude CodeClaude Haiku5.5 (medium)77*Claude CodeClaude Haiku5.5 (high)76*Claude CodeClaude Sonnet5.5 (medium)75*Terminus-2Claude Sonnet4.6(adaptive,75Terminus-2Gemini 3.1Pro (preview)(high)74*Terminus-2DeepSeek V4Flash (max)74Terminus-2DeepSeek V4Pro (high)74Claude CodeClaude Sonnet5.5 (xhigh)74*Terminus-2GLM-5.2 (max)74Terminus-2Kimi K2.7Code (alwayson)72*Terminus-2GLM-5.2(high)72Terminus-2Kimi K2.6(always on)72*Terminus-2GLM-5.3 (low)71Terminus-2DeepSeek V4Flash (high)70*Terminus-2GLM-5.2(none)68*Terminus-2GLM-5.3 (max,default)66*Terminus-2GLM-5.3-Flash(low)65Terminus-2Gemini 3.1Pro (preview)(low)62Terminus-2Gemma 4 26BA4B (thinkingon)59*Terminus-2Gemini 3.1Flash-Lite(high)56*Terminus-2GLM-5.3-Flash(max)52Terminus-2gpt-oss-120b(high)49Terminus-2Gemini 3.1Flash-Lite(low)40Terminus-2GLM-4.7-Flash(thinking on)37Terminus-2gpt-oss-120b(low)34*Terminus-2Gemini 3.1Flash-Lite(minimal)62*Terminus-2Qwen 3.8 27B(xhigh)0*Terminus-2Mistral Small3.1 24B(Non-reasoning)
Data table: Geospatial Agent Index
Geospatial Agent Index
ModelCreatorGeospatial Agent IndexRange while attempts are pendingBenchmarks coveredCoverage (attempts)
Claude Code – Claude Opus 5.5 (xhigh)Anthropic82 (attempts per task: 1)No pending judgements4 of 4290 of 290 planned attempts
Claude Code – Claude Opus 5.5 (high)Anthropic82* (attempts per task: 2.5)63 to 844 of 4531 of 870 planned attempts
Claude Code – Claude Sonnet 5.5 (high)Anthropic81* (attempts per task: 2)77 to 814 of 4556 of 870 planned attempts
Claude Code – Claude Opus 5.5 (medium)Anthropic80* (attempts per task: 2.7)58 to 834 of 4573 of 870 planned attempts
Codex – GPT-6.1 Sol (high)OpenAI80 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Codex – GPT-6.1 Sol (medium)OpenAI79* (attempts per task: 2.6)No pending judgements; coverage incomplete4 of 4749 of 870 planned attempts (6 excluded)
Codex – GPT-6.1 Sol (xhigh)OpenAI78* (attempts per task: 2.6)69 to 784 of 4668 of 870 planned attempts
Claude Code – Claude Haiku 5.5 (xhigh)Anthropic78 (attempts per task: 1)No pending judgements4 of 4290 of 290 planned attempts
Codex – GPT-6 Luna (xhigh)OpenAI77* (attempts per task: 2.4)69 to 784 of 4627 of 870 planned attempts
Claude Code – Claude Haiku 5.5 (medium)Anthropic77* (attempts per task: 2)74 to 784 of 4538 of 870 planned attempts
Claude Code – Claude Haiku 5.5 (high)Anthropic77* (attempts per task: 2.3)61 to 794 of 4526 of 870 planned attempts
Claude Code – Claude Sonnet 5.5 (medium)Anthropic76* (attempts per task: 2.2)71 to 764 of 4580 of 870 planned attempts
Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high)Anthropic75* (attempts per task: 1.2)74 to 754 of 4335 of 870 planned attempts (138 excluded)
Terminus-2 – Gemini 3.1 Pro (preview) (high)Google75 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – DeepSeek V4 Flash (max)DeepSeek74* (attempts per task: 1.2)73 to 754 of 4344 of 870 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek74 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Claude Code – Claude Sonnet 5.5 (xhigh)Anthropic74 (attempts per task: 1)No pending judgements4 of 4290 of 290 planned attempts
Terminus-2 – GLM-5.2 (max)Z.ai74* (attempts per task: 1.2)74 to 744 of 4338 of 870 planned attempts (1 excluded)
Terminus-2 – Kimi K2.7 Code (always on)Moonshot AI74 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – GLM-5.2 (high)Z.ai72* (attempts per task: 1.1)71 to 724 of 4310 of 870 planned attempts (1 excluded)
Terminus-2 – Kimi K2.6 (always on)Moonshot AI72 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – GLM-5.3 (low)Z.ai72* (attempts per task: 1)70 to 724 of 4285 of 290 planned attempts
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek71 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – GLM-5.2 (none)Z.ai70* (attempts per task: 1.6)68 to 704 of 4446 of 870 planned attempts
Terminus-2 – GLM-5.3 (max, default)Z.ai68* (attempts per task: 1.2)67 to 684 of 4335 of 870 planned attempts
Terminus-2 – GLM-5.3-Flash (low)Z.ai66* (attempts per task: 1.4)61 to 674 of 4380 of 870 planned attempts (4 excluded)
Terminus-2 – Gemini 3.1 Pro (preview) (low)Google65 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts (22 excluded)
Terminus-2 – Gemma 4 26B A4B (thinking on)Google62 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – Gemini 3.1 Flash-Lite (high)Google59* (attempts per task: 3)59 to 594 of 4864 of 870 planned attempts (6 excluded)
Terminus-2 – GLM-5.3-Flash (max)Z.ai56* (attempts per task: 1.2)55 to 564 of 4342 of 870 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI52 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – Gemini 3.1 Flash-Lite (low)Google49 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – GLM-4.7-Flash (thinking on)Z.ai40 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – gpt-oss-120b (low)OpenAI37 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – Gemini 3.1 Flash-Lite (minimal)Google34* (attempts per task: 3)No pending judgements; coverage incomplete4 of 4869 of 870 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)62* (attempts per task: 1 on 101 of 290 tasks so far)No pending judgements; coverage incomplete3 of 4101 of 870 planned attempts
Terminus-2 – Mistral Small 3.1 24B (Non-reasoning)Mistral AI0* (attempts per task: 1 on 9 of 290 tasks so far)No pending judgements; coverage incomplete1 of 49 of 351 planned attempts
Mistral AI models
ModelGeospatial Agent IndexGeoAgentBenchGeoBenchXEarth-BenchGeoAnalystBenchCost per taskTime per taskContext window
Terminus-2 – Mistral Small 3.1 24B (Non-reasoning)0* (attempts per task: 1 on 9 of 290 tasks so far)No dataNo data0No data$0.24715.6 min128,000