Axis Spatial

Terminus-2 – Gemini 3.8 Flash (low): geospatial agent results

Terminus-2 – Gemini 3.8 Flash (low) by Google, run under Terminus-2: its Geospatial Agent Index, score on each benchmark and each kind of geospatial work, cost, time and token use.

Summary

Geospatial Agent Index
64* (attempts per task: 2.9) (rank 22 of 37)
Cost per task
$0.115
Time per task
2.7 min
Tokens per attempt
314k
Turns per attempt
17.9
Attempts decided
833

Score by benchmark

Terminus-2 – Gemini 3.8 Flash (low): score by benchmark

Average pass@1, 0 to 100 · Higher is better

Terminus-2 – Gemini 3.8 Flash (low): score by benchmarkGeoAgentBench: 96 (partial coverage); GeoBenchX: 61; Earth-Bench: 37 (partial coverage); GeoAnalystBench: 63025507510096GeoAgentBench61GeoBenchX37Earth-Bench63GeoAnalystBench
Data table: Terminus-2 – Gemini 3.8 Flash (low): score by benchmark
Terminus-2 – Gemini 3.8 Flash (low): by benchmark
BenchmarkScoreTasks with a decided attemptCost per taskTime per task
GeoAgentBench9650 of 50$0.0521.5 min
GeoBenchX61173 of 173$0.1102.5 min
Earth-Bench3736 of 48$0.2265.0 min
GeoAnalystBench6319 of 19$0.1233.7 min

Score by task format and kind of work

Terminus-2 – Gemini 3.8 Flash (low): score by task format

Average pass@1 over the tasks of each task format, 0 to 100 · Higher is better

Terminus-2 – Gemini 3.8 Flash (low): score by task formatMulti-step analysis: 62 (partial coverage); Single-answer questions: 39 (partial coverage); Recognising infeasible requests: 83025507510062Multi-stepanalysis39Single-answerquestions83Recognisinginfeasiblerequests
How to read this chart

What this shows. The same attempts as the index, grouped by task format. Groups with few tasks are less certain. About these breakdowns.

Data table: Terminus-2 – Gemini 3.8 Flash (low): score by task format
Terminus-2 – Gemini 3.8 Flash (low): by task format
Task formatScoreTasks with a decided attempt
Multi-step analysis62151 of 151
Single-answer questions3948 of 159
Recognising infeasible requests8379 of 79

Terminus-2 – Gemini 3.8 Flash (low): score by kind of work

Average pass@1 over the tasks of each kind of work, 0 to 100 · Higher is better

Terminus-2 – Gemini 3.8 Flash (low): score by kind of workRaster and remote sensing: 62 (partial coverage); Vector and overlay: 63; Networks and routing: 100; Climate and time series: 51 (partial coverage); Spatial statistics and interpolation: 43; Mapping and cartography: 53; Recognising infeasible requests: 83025507510062Raster andremotesensing63Vector andoverlay100Networks androuting51Climate andtime series43Spatialstatisticsand53Mapping andcartography83Recognisinginfeasiblerequests
How to read this chart

What this shows. The same attempts as the index, grouped by kind of work. Groups with few tasks are less certain. About these breakdowns.

Data table: Terminus-2 – Gemini 3.8 Flash (low): score by kind of work
Terminus-2 – Gemini 3.8 Flash (low): by kind of work
Kind of workScoreTasks with a decided attempt
Raster and remote sensing6268 of 123
Vector and overlay6318 of 18
Networks and routing1005 of 5
Climate and time series5119 of 75
Spatial statistics and interpolation4324 of 24
Mapping and cartography5365 of 65
Recognising infeasible requests8379 of 79

Terminus-2 – Gemini 3.8 Flash (low): score by data type

Average pass@1 over the tasks of each data type, 0 to 100 · Higher is better

Terminus-2 – Gemini 3.8 Flash (low): score by data typeOptical multispectral: 30 (partial coverage); Elevation or terrain: 98; Climate or gridded time series: 55 (partial coverage); Other gridded data: 83 (partial coverage); Vector only: 55; Tabular only: 50025507510030Opticalmultispectral98Elevation orterrain55Climate orgridded timeseries83Other griddeddata55Vector only50Tabular only
How to read this chart

What this shows. The same attempts as the index, grouped by data type. Groups with few tasks are less certain. About these breakdowns.

Data table: Terminus-2 – Gemini 3.8 Flash (low): score by data type
Terminus-2 – Gemini 3.8 Flash (low): by data type
Data typeScoreTasks with a decided attempt
Optical multispectral3033 of 117
Elevation or terrain9816 of 16
Climate or gridded time series5517 of 18
Other gridded data8323 of 61
Vector only55108 of 108
Tabular only502 of 2

Terminus-2 – Gemini 3.8 Flash (low): score by task length

Average pass@1 over the tasks of each task length, 0 to 100 · Higher is better

Terminus-2 – Gemini 3.8 Flash (low): score by task lengthShort tasks: 92 (partial coverage); Medium tasks: 63 (partial coverage); Long tasks: 37 (partial coverage)025507510092Short tasks63Medium tasks37Long tasks
How to read this chart

What this shows. The same attempts as the index, grouped by task length. Groups with few tasks are less certain. About these breakdowns.

Data table: Terminus-2 – Gemini 3.8 Flash (low): score by task length
Terminus-2 – Gemini 3.8 Flash (low): by task length
Task lengthScoreTasks with a decided attempt
Short tasks9287 of 129
Medium tasks63107 of 131
Long tasks3784 of 129

Compared with other models

Geospatial Agent Index

Equal-weight average of the evaluation scores, 0 to 100 · Higher is better

  • Partial coverage
Geospatial Agent IndexTerminus-2 – DeepSeek V4 Pro (high): 72* (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 72* (partial coverage); Terminus-2 – GLM-5.2 (high): 72* (partial coverage); Terminus-2 – DeepSeek V4 Flash (max): 71* (partial coverage); Codex – GPT-6.1 Sol (medium): 70* (partial coverage); Codex – GPT-6.1 Sol (high): 69* (partial coverage); Terminus-2 – GLM-5.2 (max): 69* (partial coverage); Claude Code – Claude Opus 5.5 (high): 69* (partial coverage); Codex – GPT-6 Luna (xhigh): 69* (partial coverage); Codex – GPT-6.1 Sol (xhigh): 68* (partial coverage); Claude Code – Claude Opus 5.5 (xhigh): 68* (partial coverage); Claude Code – Claude Haiku 5.5 (xhigh): 68* (partial coverage); Terminus-2 – Kimi K2.7 Code (always on): 68* (partial coverage); Claude Code – Claude Sonnet 5.5 (high): 68* (partial coverage); Claude Code – Claude Opus 5.5 (medium): 67* (partial coverage); Terminus-2 – Kimi K2.6 (always on): 67* (partial coverage); Claude Code – Claude Haiku 5.5 (high): 66* (partial coverage); Claude Code – Claude Haiku 5.5 (medium): 65* (partial coverage); Claude Code – Claude Sonnet 5.5 (medium): 65* (partial coverage); Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high): 65* (partial coverage); Terminus-2 – Gemini 3.1 Pro (preview) (high): 65* (partial coverage); Terminus-2 – Gemini 3.8 Flash (low): 64* (partial coverage); Claude Code – Claude Sonnet 5.5 (xhigh): 64* (partial coverage); Terminus-2 – GLM-5.3 (low): 64* (partial coverage); Terminus-2 – GLM-5.3 (max, default): 64* (partial coverage); Terminus-2 – GLM-5.2 (none): 64* (partial coverage); Terminus-2 – Gemini 3.1 Pro (preview) (low): 63* (partial coverage); Terminus-2 – GLM-5.3-Flash (low): 62* (partial coverage); Terminus-2 – GLM-5.3-Flash (max): 56* (partial coverage); Terminus-2 – Gemma 4 26B A4B (thinking on): 56* (partial coverage); Terminus-2 – Gemini 3.1 Flash-Lite (low): 51* (partial coverage); Terminus-2 – Gemini 3.1 Flash-Lite (high): 51* (partial coverage); Terminus-2 – gpt-oss-120b (high): 48* (partial coverage); Terminus-2 – Gemini 3.8 Flash (high): 39* (partial coverage); Terminus-2 – gpt-oss-120b (low): 38* (partial coverage); Terminus-2 – Gemini 3.1 Flash-Lite (minimal): 33* (partial coverage); Terminus-2 – GLM-4.7-Flash (thinking on): 33* (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 55* (partial coverage); Terminus-2 – Mistral Small 3.1 24B (Non-reasoning): 0* (partial coverage)025507510072*Terminus-2DeepSeek V4Pro (high)72*Terminus-2DeepSeek V4Flash (high)72*Terminus-2GLM-5.2(high)71*Terminus-2DeepSeek V4Flash (max)70*CodexGPT-6.1 Sol(medium)69*CodexGPT-6.1 Sol(high)69*Terminus-2GLM-5.2 (max)69*Claude CodeClaude Opus5.5 (high)69*CodexGPT-6 Luna(xhigh)68*CodexGPT-6.1 Sol(xhigh)68*Claude CodeClaude Opus5.5 (xhigh)68*Claude CodeClaude Haiku5.5 (xhigh)68*Terminus-2Kimi K2.7Code (alwayson)68*Claude CodeClaude Sonnet5.5 (high)67*Claude CodeClaude Opus5.5 (medium)67*Terminus-2Kimi K2.6(always on)66*Claude CodeClaude Haiku5.5 (high)65*Claude CodeClaude Haiku5.5 (medium)65*Claude CodeClaude Sonnet5.5 (medium)65*Terminus-2Claude Sonnet4.6(adaptive,65*Terminus-2Gemini 3.1Pro (preview)(high)64*Terminus-2Gemini 3.8Flash (low)64*Claude CodeClaude Sonnet5.5 (xhigh)64*Terminus-2GLM-5.3 (low)64*Terminus-2GLM-5.3 (max,default)64*Terminus-2GLM-5.2(none)63*Terminus-2Gemini 3.1Pro (preview)(low)62*Terminus-2GLM-5.3-Flash(low)56*Terminus-2GLM-5.3-Flash(max)56*Terminus-2Gemma 4 26BA4B (thinkingon)51*Terminus-2Gemini 3.1Flash-Lite(low)51*Terminus-2Gemini 3.1Flash-Lite(high)48*Terminus-2gpt-oss-120b(high)39*Terminus-2Gemini 3.8Flash (high)38*Terminus-2gpt-oss-120b(low)33*Terminus-2Gemini 3.1Flash-Lite(minimal)33*Terminus-2GLM-4.7-Flash(thinking on)55*Terminus-2Qwen 3.8 27B(xhigh)0*Terminus-2Mistral Small3.1 24B(Non-reasoning)
How to read this chart

What this metric means. The equal-weight average of the 4 evaluation scores. Each evaluation score is the average pass@1 over that benchmark's tasks, from 0 to 100.

Not yet ranked. Entries without a score on every benchmark come after the ranked entries and are not ranked: their index averages only the benchmarks they have.

Partial coverage. A hatched bar, and an asterisk in tables, marks an entry that does not yet have results for every planned attempt. While attempts await review, the table gives the range the score could still reach.

Attempts per task. Each task's score is the average of its recorded attempts (up to 3). Some entries and tracks currently have 1 attempt per task; further attempts will be added and averaged in. Each entry shows its attempts per task.

Data table: Geospatial Agent Index
Geospatial Agent Index
ModelCreatorGeospatial Agent IndexRange while attempts are pendingBenchmarks coveredCoverage (attempts)
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek72* (attempts per task: 3.8)30 to 844 of 4544 of 1,167 planned attempts
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek72* (attempts per task: 3.2)30 to 844 of 4368 of 1,167 planned attempts (21 excluded)
Terminus-2 – GLM-5.2 (high)Z.ai72* (attempts per task: 1.8)68 to 734 of 4484 of 1,167 planned attempts (3 excluded)
Terminus-2 – DeepSeek V4 Flash (max)DeepSeek71* (attempts per task: 1.9)65 to 734 of 4464 of 1,167 planned attempts (19 excluded)
Codex – GPT-6.1 Sol (medium)OpenAI70* (attempts per task: 3.3)67 to 804 of 4870 of 1,167 planned attempts (1 excluded)
Codex – GPT-6.1 Sol (high)OpenAI69* (attempts per task: 3.3)66 to 784 of 4877 of 1,167 planned attempts
Terminus-2 – GLM-5.2 (max)Z.ai69* (attempts per task: 2)65 to 714 of 4508 of 1,167 planned attempts (3 excluded)
Claude Code – Claude Opus 5.5 (high)Anthropic69* (attempts per task: 3.9)66 to 804 of 4921 of 1,167 planned attempts
Codex – GPT-6 Luna (xhigh)OpenAI69* (attempts per task: 3.3)34 to 864 of 4444 of 1,167 planned attempts
Codex – GPT-6.1 Sol (xhigh)OpenAI68* (attempts per task: 3.3)66 to 784 of 4871 of 1,167 planned attempts
Claude Code – Claude Opus 5.5 (xhigh)Anthropic68* (attempts per task: 1.3)65 to 784 of 4308 of 683 planned attempts
Claude Code – Claude Haiku 5.5 (xhigh)Anthropic68* (attempts per task: 1.3)66 to 784 of 4319 of 683 planned attempts
Terminus-2 – Kimi K2.7 Code (always on)Moonshot AI68* (attempts per task: 3.7)28 to 834 of 4496 of 1,167 planned attempts
Claude Code – Claude Sonnet 5.5 (high)Anthropic68* (attempts per task: 4)66 to 794 of 4949 of 1,167 planned attempts
Claude Code – Claude Opus 5.5 (medium)Anthropic67* (attempts per task: 4)66 to 784 of 4942 of 1,167 planned attempts
Terminus-2 – Kimi K2.6 (always on)Moonshot AI67* (attempts per task: 3.8)27 to 824 of 4532 of 1,167 planned attempts
Claude Code – Claude Haiku 5.5 (high)Anthropic66* (attempts per task: 3.8)64 to 764 of 4937 of 1,167 planned attempts
Claude Code – Claude Haiku 5.5 (medium)Anthropic65* (attempts per task: 4)47 to 804 of 4745 of 1,167 planned attempts
Claude Code – Claude Sonnet 5.5 (medium)Anthropic65* (attempts per task: 4)65 to 774 of 4951 of 1,167 planned attempts
Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high)Anthropic65* (attempts per task: 1.1)62 to 664 of 4295 of 1,167 planned attempts (149 excluded)
Terminus-2 – Gemini 3.1 Pro (preview) (high)Google65* (attempts per task: 2.9)No pending judgements; coverage incomplete4 of 4834 of 1,167 planned attempts
Terminus-2 – Gemini 3.8 Flash (low)Google64* (attempts per task: 2.9)No pending judgements; coverage incomplete4 of 4833 of 1,167 planned attempts (74 excluded)
Claude Code – Claude Sonnet 5.5 (xhigh)Anthropic64* (attempts per task: 1.3)62 to 744 of 4318 of 683 planned attempts
Terminus-2 – GLM-5.3 (low)Z.ai64* (attempts per task: 1.8)63 to 644 of 4513 of 1,167 planned attempts
Terminus-2 – GLM-5.3 (max, default)Z.ai64* (attempts per task: 1.8)60 to 654 of 4463 of 1,167 planned attempts
Terminus-2 – GLM-5.2 (none)Z.ai64* (attempts per task: 2.9)50 to 694 of 4636 of 1,167 planned attempts (2 excluded)
Terminus-2 – Gemini 3.1 Pro (preview) (low)Google63* (attempts per task: 2.9)21 to 884 of 4278 of 1,167 planned attempts (21 excluded)
Terminus-2 – GLM-5.3-Flash (low)Z.ai62* (attempts per task: 1.6)45 to 724 of 4298 of 1,167 planned attempts (10 excluded)
Terminus-2 – GLM-5.3-Flash (max)Z.ai56* (attempts per task: 1.9)51 to 594 of 4493 of 1,167 planned attempts
Terminus-2 – Gemma 4 26B A4B (thinking on)Google56* (attempts per task: 3.7)22 to 774 of 4516 of 1,167 planned attempts
Terminus-2 – Gemini 3.1 Flash-Lite (low)Google51* (attempts per task: 2.9)17 to 844 of 4278 of 1,167 planned attempts
Terminus-2 – Gemini 3.1 Flash-Lite (high)Google51* (attempts per task: 2.9)17 to 844 of 4272 of 1,167 planned attempts (6 excluded)
Terminus-2 – gpt-oss-120b (high)OpenAI48* (attempts per task: 4)20 to 744 of 4611 of 1,167 planned attempts
Terminus-2 – Gemini 3.8 Flash (high)Google39* (attempts per task: 2.9)13 to 804 of 4278 of 1,167 planned attempts (1 excluded)
Terminus-2 – gpt-oss-120b (low)OpenAI38* (attempts per task: 4)15 to 694 of 4611 of 1,167 planned attempts
Terminus-2 – Gemini 3.1 Flash-Lite (minimal)Google33* (attempts per task: 2.9)11 to 784 of 4278 of 1,167 planned attempts
Terminus-2 – GLM-4.7-Flash (thinking on)Z.ai33* (attempts per task: 3.6)14 to 694 of 4483 of 1,167 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)55* (attempts per task: 1 on 89 of 290 tasks so far)No pending judgements; coverage incomplete3 of 489 of 1,167 planned attempts
Terminus-2 – Mistral Small 3.1 24B (Non-reasoning)Mistral AI0* (attempts per task: 1 on 9 of 290 tasks so far)No pending judgements; coverage incomplete1 of 49 of 648 planned attempts (4 excluded)

Geospatial Agent Index vs. cost per task

Higher and further left is better

  • Partial coverage
  • Most attractive quadrant
  • Pareto line
Geospatial Agent Index vs. cost per taskTerminus-2 – Claude Sonnet 4.6 (adaptive, effort high): Geospatial Agent Index 64.94, cost per task (usd) $0.252; Terminus-2 – DeepSeek V4 Flash (max): Geospatial Agent Index 71.07, cost per task (usd) $0.067; Terminus-2 – DeepSeek V4 Flash (high): Geospatial Agent Index 72.04, cost per task (usd) $0.070; Terminus-2 – DeepSeek V4 Pro (high): Geospatial Agent Index 72.29, cost per task (usd) $0.365; Terminus-2 – Gemini 3.1 Flash-Lite (low): Geospatial Agent Index 51.22, cost per task (usd) $0.011; Terminus-2 – Gemini 3.1 Flash-Lite (minimal): Geospatial Agent Index 33.25, cost per task (usd) $0.0054; Terminus-2 – Gemini 3.1 Flash-Lite (high): Geospatial Agent Index 51.2, cost per task (usd) $0.020; Terminus-2 – Gemini 3.1 Pro (preview) (low): Geospatial Agent Index 62.71, cost per task (usd) $0.060; Terminus-2 – Gemini 3.1 Pro (preview) (high): Geospatial Agent Index 64.89, cost per task (usd) $0.237; Terminus-2 – Gemini 3.8 Flash (high): Geospatial Agent Index 39.45, cost per task (usd) $0.308; Terminus-2 – Gemini 3.8 Flash (low): Geospatial Agent Index 64.22, cost per task (usd) $0.115; Terminus-2 – Gemma 4 26B A4B (thinking on): Geospatial Agent Index 55.66, cost per task (usd) $0.017; Terminus-2 – GLM-4.7-Flash (thinking on): Geospatial Agent Index 33.18, cost per task (usd) $0.033; Terminus-2 – GLM-5.2 (high): Geospatial Agent Index 71.6, cost per task (usd) $0.255; Terminus-2 – GLM-5.2 (none): Geospatial Agent Index 63.56, cost per task (usd) $0.104; Terminus-2 – GLM-5.2 (max): Geospatial Agent Index 69, cost per task (usd) $0.361; Terminus-2 – GLM-5.3-Flash (low): Geospatial Agent Index 61.81, cost per task (usd) $0.027; Terminus-2 – GLM-5.3-Flash (max): Geospatial Agent Index 56.05, cost per task (usd) $0.050; Terminus-2 – GLM-5.3 (low): Geospatial Agent Index 63.9, cost per task (usd) $0.326; Terminus-2 – GLM-5.3 (max, default): Geospatial Agent Index 63.66, cost per task (usd) $0.434; Terminus-2 – gpt-oss-120b (low): Geospatial Agent Index 37.77, cost per task (usd) $0.0080; Terminus-2 – gpt-oss-120b (high): Geospatial Agent Index 48.13, cost per task (usd) $0.031; Terminus-2 – Kimi K2.6 (always on): Geospatial Agent Index 66.72, cost per task (usd) $0.093; Terminus-2 – Kimi K2.7 Code (always on): Geospatial Agent Index 68, cost per task (usd) $0.142; Claude Code – Claude Opus 5.5 (medium): Geospatial Agent Index 67.32, cost per task (usd) $0.139; Claude Code – Claude Opus 5.5 (high): Geospatial Agent Index 68.94, cost per task (usd) $0.177; Claude Code – Claude Sonnet 5.5 (medium): Geospatial Agent Index 65.25, cost per task (usd) $0.052; Claude Code – Claude Sonnet 5.5 (high): Geospatial Agent Index 67.93, cost per task (usd) $0.067; Claude Code – Claude Haiku 5.5 (medium): Geospatial Agent Index 65.36, cost per task (usd) $0.0074; Claude Code – Claude Haiku 5.5 (high): Geospatial Agent Index 65.58, cost per task (usd) $0.011; Claude Code – Claude Opus 5.5 (xhigh): Geospatial Agent Index 68.39, cost per task (usd) $0.398; Claude Code – Claude Sonnet 5.5 (xhigh): Geospatial Agent Index 63.95, cost per task (usd) $0.173; Claude Code – Claude Haiku 5.5 (xhigh): Geospatial Agent Index 68.3, cost per task (usd) $0.024; Codex – GPT-6.1 Sol (medium): Geospatial Agent Index 70.25, cost per task (usd) $0.060; Codex – GPT-6.1 Sol (high): Geospatial Agent Index 69.19, cost per task (usd) $0.080; Codex – GPT-6.1 Sol (xhigh): Geospatial Agent Index 68.46, cost per task (usd) $0.124; Codex – GPT-6 Luna (xhigh): Geospatial Agent Index 68.68, cost per task (usd) $0.0095Most attractive quadrant253035404550556065707580$0.002$0.005$0.01$0.02$0.05$0.10$0.20$0.50Cost per task (USD) (log scale)Geospatial Agent IndexTerminus-2 – DeepSeek V4 Pro (high)Terminus-2 – DeepSeek V4 Flash (high)Terminus-2 – GLM-5.2 (high)Terminus-2 – DeepSeek V4 Flash (max)Codex – GPT-6.1 Sol (medium)Codex – GPT-6.1 Sol (high)Terminus-2 – GLM-5.2 (max)Claude Code – Claude Opus 5.5 (high)Codex – GPT-6 Luna (xhigh)Codex – GPT-6.1 Sol (xhigh)Claude Code – Claude Opus 5.5 (xhigh)Claude Code – Claude Haiku 5.5 (xhigh)Terminus-2 – Kimi K2.7 Code (always on)Claude Code – Claude Sonnet 5.5 (high)Claude Code – Claude Opus 5.5 (medium)Terminus-2 – Kimi K2.6 (always on)Claude Code – Claude Haiku 5.5 (high)Claude Code – Claude Haiku 5.5 (medium)Claude Code – Claude Sonnet 5.5 (medium)Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high)Terminus-2 – Gemini 3.1 Pro (preview) (high)Terminus-2 – Gemini 3.8 Flash (low)Claude Code – Claude Sonnet 5.5 (xhigh)Terminus-2 – GLM-5.3 (low)Terminus-2 – GLM-5.3 (max, default)Terminus-2 – GLM-5.2 (none)Terminus-2 – Gemini 3.1 Pro (preview) (low)Terminus-2 – GLM-5.3-Flash (low)Terminus-2 – GLM-5.3-Flash (max)Terminus-2 – Gemma 4 26B A4B (thinking on)Terminus-2 – Gemini 3.1 Flash-Lite (low)Terminus-2 – Gemini 3.1 Flash-Lite (high)Terminus-2 – gpt-oss-120b (high)Terminus-2 – Gemini 3.8 Flash (high)Terminus-2 – gpt-oss-120b (low)Terminus-2 – Gemini 3.1 Flash-Lite (minimal)Terminus-2 – GLM-4.7-Flash (thinking on)
How to read this chart

Not yet ranked. Not plotted, so the axes fit the ranked entries; listed in the table: Terminus-2 – Mistral Small 3.1 24B (Non-reasoning), Terminus-2 – Qwen 3.8 27B (xhigh).

What cost is measuring. What one attempt at a task would cost if paid token by token at the provider's published API prices, averaged over attempts: ordinary input tokens, input tokens served from the provider's cache at the lower cached price, tokens written to the cache where the provider charges for that, and output tokens (thinking included). Cost assumes each harness uses the model's pay-per-token API, priced at public list rates; sandboxes and our own engineering time are not included. Failed attempts count too.

How to read this chart. Each point is one model configuration. Further left means a lower average cost per task (log scale); higher means a higher Geospatial Agent Index. Both axes fit the models shown, so the score axis does not start at zero. The shaded most attractive quadrant is the top-left quarter of the plot, split at the middle of each axis: a reading aid that moves with the axes when you change the models shown, not a fixed bar. The dotted Pareto line joins the shown models that no cheaper shown model beats on score.

Data table: Geospatial Agent Index vs. cost per task
Geospatial Agent Index vs. cost per task
ModelGeospatial Agent IndexCost per task (USD)On the Pareto line
Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high)65* (attempts per task: 1.1)$0.252No
Terminus-2 – DeepSeek V4 Flash (max)71* (attempts per task: 1.9)$0.067Yes
Terminus-2 – DeepSeek V4 Flash (high)72* (attempts per task: 3.2)$0.070Yes
Terminus-2 – DeepSeek V4 Pro (high)72* (attempts per task: 3.8)$0.365Yes
Terminus-2 – Gemini 3.1 Flash-Lite (low)51* (attempts per task: 2.9)$0.011No
Terminus-2 – Gemini 3.1 Flash-Lite (minimal)33* (attempts per task: 2.9)$0.0054Yes
Terminus-2 – Gemini 3.1 Flash-Lite (high)51* (attempts per task: 2.9)$0.020No
Terminus-2 – Gemini 3.1 Pro (preview) (low)63* (attempts per task: 2.9)$0.060No
Terminus-2 – Gemini 3.1 Pro (preview) (high)65* (attempts per task: 2.9)$0.237No
Terminus-2 – Gemini 3.8 Flash (high)39* (attempts per task: 2.9)$0.308No
Terminus-2 – Gemini 3.8 Flash (low)64* (attempts per task: 2.9)$0.115No
Terminus-2 – Gemma 4 26B A4B (thinking on)56* (attempts per task: 3.7)$0.017No
Terminus-2 – GLM-4.7-Flash (thinking on)33* (attempts per task: 3.6)$0.033No
Terminus-2 – GLM-5.2 (high)72* (attempts per task: 1.8)$0.255No
Terminus-2 – GLM-5.2 (none)64* (attempts per task: 2.9)$0.104No
Terminus-2 – GLM-5.2 (max)69* (attempts per task: 2)$0.361No
Terminus-2 – GLM-5.3-Flash (low)62* (attempts per task: 1.6)$0.027No
Terminus-2 – GLM-5.3-Flash (max)56* (attempts per task: 1.9)$0.050No
Terminus-2 – GLM-5.3 (low)64* (attempts per task: 1.8)$0.326No
Terminus-2 – GLM-5.3 (max, default)64* (attempts per task: 1.8)$0.434No
Terminus-2 – gpt-oss-120b (low)38* (attempts per task: 4)$0.0080No
Terminus-2 – gpt-oss-120b (high)48* (attempts per task: 4)$0.031No
Terminus-2 – Kimi K2.6 (always on)67* (attempts per task: 3.8)$0.093No
Terminus-2 – Kimi K2.7 Code (always on)68* (attempts per task: 3.7)$0.142No
Terminus-2 – Mistral Small 3.1 24B (Non-reasoning)0* (attempts per task: 1 on 9 of 290 tasks so far)$0.215Not plotted: not yet ranked
Terminus-2 – Qwen 3.8 27B (xhigh)55* (attempts per task: 1 on 89 of 290 tasks so far)$0.063Not plotted: not yet ranked
Claude Code – Claude Opus 5.5 (medium)67* (attempts per task: 4)$0.139No
Claude Code – Claude Opus 5.5 (high)69* (attempts per task: 3.9)$0.177No
Claude Code – Claude Sonnet 5.5 (medium)65* (attempts per task: 4)$0.052No
Claude Code – Claude Sonnet 5.5 (high)68* (attempts per task: 4)$0.067No
Claude Code – Claude Haiku 5.5 (medium)65* (attempts per task: 4)$0.0074Yes
Claude Code – Claude Haiku 5.5 (high)66* (attempts per task: 3.8)$0.011No
Claude Code – Claude Opus 5.5 (xhigh)68* (attempts per task: 1.3)$0.398No
Claude Code – Claude Sonnet 5.5 (xhigh)64* (attempts per task: 1.3)$0.173No
Claude Code – Claude Haiku 5.5 (xhigh)68* (attempts per task: 1.3)$0.024No
Codex – GPT-6.1 Sol (medium)70* (attempts per task: 3.3)$0.060Yes
Codex – GPT-6.1 Sol (high)69* (attempts per task: 3.3)$0.080No
Codex – GPT-6.1 Sol (xhigh)68* (attempts per task: 3.3)$0.124No
Codex – GPT-6 Luna (xhigh)69* (attempts per task: 3.3)$0.0095Yes

Head-to-head comparisons

Specification and settings

Developer
Google
Context window
1,048,576 tokens
Image input
Yes
Reasoning setting
low
Temperature
1.0
Maximum output tokens
65,536
Input price per 1M tokens
$0.75
Cached input price per 1M tokens
$0.075
Output price per 1M tokens
$3.75

Why this model was chosen: model selection.