Axis Spatial

Anthropic models: geospatial agent results

Geospatial Agent Benchmarks tests 10 model configurations from Anthropic.

Highlights

Results

Geospatial Agent Index

Equal-weight average of the evaluation scores, 0 to 100 · Higher is better

  • Partial coverage
Geospatial Agent IndexClaude Code – Claude Opus 5.5 (high): 83* (partial coverage); Claude Code – Claude Opus 5.5 (xhigh): 82; Claude Code – Claude Opus 5.5 (medium): 81* (partial coverage); Claude Code – Claude Sonnet 5.5 (high): 81* (partial coverage); Codex – GPT-6.1 Sol (high): 80* (partial coverage); Codex – GPT-6.1 Sol (medium): 79* (partial coverage); Codex – GPT-6.1 Sol (xhigh): 78* (partial coverage); Claude Code – Claude Haiku 5.5 (xhigh): 78; Codex – GPT-6 Luna (xhigh): 77* (partial coverage); Claude Code – Claude Haiku 5.5 (medium): 77* (partial coverage); Claude Code – Claude Sonnet 5.5 (medium): 77* (partial coverage); Claude Code – Claude Haiku 5.5 (high): 75* (partial coverage); Terminus-2 – GLM-5.2 (max): 75* (partial coverage); Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high): 74* (partial coverage); Terminus-2 – DeepSeek V4 Flash (max): 74* (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 74; Claude Code – Claude Sonnet 5.5 (xhigh): 74; Terminus-2 – Kimi K2.7 Code (always on): 74; Terminus-2 – Kimi K2.6 (always on): 72; Terminus-2 – GLM-5.2 (high): 71* (partial coverage); Terminus-2 – GLM-5.3 (low): 71* (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 70* (partial coverage); Terminus-2 – GLM-5.2 (none): 69* (partial coverage); Terminus-2 – GLM-5.3 (max, default): 67* (partial coverage); Terminus-2 – GLM-5.3-Flash (low): 66* (partial coverage); Terminus-2 – Gemma 4 26B A4B (thinking on): 62; Terminus-2 – GLM-5.3-Flash (max): 56* (partial coverage); Terminus-2 – gpt-oss-120b (high): 52; Terminus-2 – GLM-4.7-Flash (thinking on): 40; Terminus-2 – gpt-oss-120b (low): 37; Terminus-2 – Qwen 3.8 27B (xhigh): 62* (partial coverage)025507510083*Claude CodeClaude Opus5.5 (high)82Claude CodeClaude Opus5.5 (xhigh)81*Claude CodeClaude Opus5.5 (medium)81*Claude CodeClaude Sonnet5.5 (high)80*CodexGPT-6.1 Sol(high)79*CodexGPT-6.1 Sol(medium)78*CodexGPT-6.1 Sol(xhigh)78Claude CodeClaude Haiku5.5 (xhigh)77*CodexGPT-6 Luna(xhigh)77*Claude CodeClaude Haiku5.5 (medium)77*Claude CodeClaude Sonnet5.5 (medium)75*Claude CodeClaude Haiku5.5 (high)75*Terminus-2GLM-5.2 (max)74*Terminus-2Claude Sonnet4.6(adaptive,74*Terminus-2DeepSeek V4Flash (max)74Terminus-2DeepSeek V4Pro (high)74Claude CodeClaude Sonnet5.5 (xhigh)74Terminus-2Kimi K2.7Code (alwayson)72Terminus-2Kimi K2.6(always on)71*Terminus-2GLM-5.2(high)71*Terminus-2GLM-5.3 (low)70*Terminus-2DeepSeek V4Flash (high)69*Terminus-2GLM-5.2(none)67*Terminus-2GLM-5.3 (max,default)66*Terminus-2GLM-5.3-Flash(low)62Terminus-2Gemma 4 26BA4B (thinkingon)56*Terminus-2GLM-5.3-Flash(max)52Terminus-2gpt-oss-120b(high)40Terminus-2GLM-4.7-Flash(thinking on)37Terminus-2gpt-oss-120b(low)62*Terminus-2Qwen 3.8 27B(xhigh)
Data table: Geospatial Agent Index
Geospatial Agent Index
ModelCreatorGeospatial Agent IndexRange while attempts are pendingBenchmarks coveredCoverage (attempts)
Claude Code – Claude Opus 5.5 (high)Anthropic83* (partial: 1.7 of 3 attempts)71 to 834 of 4422 of 870 planned attempts
Claude Code – Claude Opus 5.5 (xhigh)Anthropic82 (partial: 1 of 3 attempts)No pending judgements4 of 4290 of 290 planned attempts
Claude Code – Claude Opus 5.5 (medium)Anthropic81* (partial: 1.8 of 3 attempts)70 to 814 of 4461 of 870 planned attempts
Claude Code – Claude Sonnet 5.5 (high)Anthropic81* (partial: 1.9 of 3 attempts)71 to 824 of 4479 of 870 planned attempts (236 excluded)
Codex – GPT-6.1 Sol (high)OpenAI80* (partial: 2.8 of 3 attempts)72 to 804 of 4753 of 870 planned attempts
Codex – GPT-6.1 Sol (medium)OpenAI79* (partial: 2.6 of 3 attempts)No pending judgements; coverage incomplete4 of 4749 of 870 planned attempts (6 excluded)
Codex – GPT-6.1 Sol (xhigh)OpenAI78* (partial: 2 of 3 attempts)74 to 794 of 4551 of 870 planned attempts
Claude Code – Claude Haiku 5.5 (xhigh)Anthropic78 (partial: 1 of 3 attempts)No pending judgements4 of 4290 of 290 planned attempts
Codex – GPT-6 Luna (xhigh)OpenAI77* (partial: 1.9 of 3 attempts)72 to 784 of 4512 of 870 planned attempts
Claude Code – Claude Haiku 5.5 (medium)Anthropic77* (partial: 1.8 of 3 attempts)66 to 784 of 4462 of 870 planned attempts
Claude Code – Claude Sonnet 5.5 (medium)Anthropic77* (partial: 2 of 3 attempts)67 to 774 of 4500 of 870 planned attempts (229 excluded)
Claude Code – Claude Haiku 5.5 (high)Anthropic75* (partial: 1.7 of 3 attempts)66 to 774 of 4422 of 870 planned attempts
Terminus-2 – GLM-5.2 (max)Z.ai75* (partial: 1.1 of 3 attempts)72 to 754 of 4306 of 870 planned attempts
Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high)Anthropic74* (partial: 1.1 of 3 attempts)73 to 754 of 4306 of 870 planned attempts (133 excluded)
Terminus-2 – DeepSeek V4 Flash (max)DeepSeek74* (partial: 1.1 of 3 attempts)73 to 744 of 4321 of 870 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek74No pending judgements4 of 4870 of 870 planned attempts
Claude Code – Claude Sonnet 5.5 (xhigh)Anthropic74 (partial: 1 of 3 attempts)No pending judgements4 of 4290 of 290 planned attempts
Terminus-2 – Kimi K2.7 Code (always on)Moonshot AI74No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – Kimi K2.6 (always on)Moonshot AI72No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – GLM-5.2 (high)Z.ai71* (partial: 1 of 3 attempts)65 to 724 of 4273 of 290 planned attempts
Terminus-2 – GLM-5.3 (low)Z.ai71* (partial: 0.9 of 3 attempts)61 to 724 of 4238 of 290 planned attempts
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek70*70 to 714 of 4860 of 870 planned attempts
Terminus-2 – GLM-5.2 (none)Z.ai69* (partial: 1.4 of 3 attempts)62 to 724 of 4356 of 870 planned attempts
Terminus-2 – GLM-5.3 (max, default)Z.ai67* (partial: 1.1 of 3 attempts)66 to 684 of 4310 of 870 planned attempts
Terminus-2 – GLM-5.3-Flash (low)Z.ai66* (partial: 1.1 of 3 attempts)59 to 674 of 4288 of 870 planned attempts (3 excluded)
Terminus-2 – Gemma 4 26B A4B (thinking on)Google62No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – GLM-5.3-Flash (max)Z.ai56* (partial: 1.1 of 3 attempts)54 to 564 of 4298 of 870 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI52No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – GLM-4.7-Flash (thinking on)Z.ai40No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – gpt-oss-120b (low)OpenAI37No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)62* (partial: 0.3 of 3 attempts)No pending judgements; coverage incomplete3 of 4101 of 870 planned attempts
Anthropic models
ModelGeospatial Agent IndexGeoAgentBenchGeoBenchXEarth-BenchGeoAnalystBenchCost per taskTime per taskContext window
Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high)74* (partial: 1.1 of 3 attempts)97515495$0.2783.7 min1,000,000
Claude Code – Claude Opus 5.5 (medium)81* (partial: 1.8 of 3 attempts)966662100$0.1380.7 min0
Claude Code – Claude Opus 5.5 (high)83* (partial: 1.7 of 3 attempts)966669100$0.1751.0 min0
Claude Code – Claude Sonnet 5.5 (medium)77* (partial: 2 of 3 attempts)98664597$0.0520.6 min0
Claude Code – Claude Sonnet 5.5 (high)81* (partial: 1.9 of 3 attempts)100675897$0.0620.7 min0
Claude Code – Claude Haiku 5.5 (medium)77* (partial: 1.8 of 3 attempts)98615395$0.00750.8 min0
Claude Code – Claude Haiku 5.5 (high)75* (partial: 1.7 of 3 attempts)98615189$0.0101.1 min0
Claude Code – Claude Opus 5.5 (xhigh)82 (partial: 1 of 3 attempts)986962100$0.3752.1 min0
Claude Code – Claude Sonnet 5.5 (xhigh)74 (partial: 1 of 3 attempts)94654889$0.1631.4 min0
Claude Code – Claude Haiku 5.5 (xhigh)78 (partial: 1 of 3 attempts)100655295$0.0231.9 min0