Axis Spatial

Terminus-2 – Gemini 3.1 Flash-Lite (low): geospatial agent results

Terminus-2 – Gemini 3.1 Flash-Lite (low) by Google, run under the Terminus-2 agent: its Geospatial Agent Index, score on each benchmark and each kind of geospatial work, cost, time and token use.

Summary

Geospatial Agent Index
49 (attempts per task: 3) (rank 32 of 35)
Cost per task
$0.012
Time per task
1.0 min
Tokens per attempt
37k
Turns per attempt
5.1
Attempts decided
870

Score by benchmark

Terminus-2 – Gemini 3.1 Flash-Lite (low): score by benchmark

Average pass@1, 0 to 100 · Higher is better

Terminus-2 – Gemini 3.1 Flash-Lite (low): score by benchmarkGeoAgentBench: 52; GeoBenchX: 46; Earth-Bench: 40; GeoAnalystBench: 60025507510052GeoAgentBench46GeoBenchX40Earth-Bench60GeoAnalystBench
Data table: Terminus-2 – Gemini 3.1 Flash-Lite (low): score by benchmark
Terminus-2 – Gemini 3.1 Flash-Lite (low): by benchmark
BenchmarkScoreTasks with a decided attemptCost per taskTime per task
GeoAgentBench5250 of 50$0.00830.8 min
GeoBenchX46173 of 173$0.00980.9 min
Earth-Bench4048 of 48$0.0241.9 min
GeoAnalystBench6019 of 19$0.00760.8 min

Score by task format and kind of work

Terminus-2 – Gemini 3.1 Flash-Lite (low): score by task format

Average pass@1 over the tasks of each task format, 0 to 100 · Higher is better

Terminus-2 – Gemini 3.1 Flash-Lite (low): score by task formatMulti-step analysis: 49; Single-answer questions: 41; Recognising infeasible requests: 48025507510049Multi-stepanalysis41Single-answerquestions48Recognisinginfeasiblerequests
How to read this chart

What this shows. The same attempts as the index, grouped by task format. Groups with few tasks are less certain. About these breakdowns.

Data table: Terminus-2 – Gemini 3.1 Flash-Lite (low): score by task format
Terminus-2 – Gemini 3.1 Flash-Lite (low): by task format
Task formatScoreTasks with a decided attempt
Multi-step analysis49151 of 151
Single-answer questions4160 of 60
Recognising infeasible requests4879 of 79

Terminus-2 – Gemini 3.1 Flash-Lite (low): score by kind of work

Average pass@1 over the tasks of each kind of work, 0 to 100 · Higher is better

Terminus-2 – Gemini 3.1 Flash-Lite (low): score by kind of workRaster and remote sensing: 46; Vector and overlay: 46; Networks and routing: 53; Climate and time series: 56; Spatial statistics and interpolation: 29; Mapping and cartography: 50; Recognising infeasible requests: 48025507510046Raster andremotesensing46Vector andoverlay53Networks androuting56Climate andtime series29Spatialstatisticsand50Mapping andcartography48Recognisinginfeasiblerequests
How to read this chart

What this shows. The same attempts as the index, grouped by kind of work. Groups with few tasks are less certain. About these breakdowns.

Data table: Terminus-2 – Gemini 3.1 Flash-Lite (low): score by kind of work
Terminus-2 – Gemini 3.1 Flash-Lite (low): by kind of work
Kind of workScoreTasks with a decided attempt
Raster and remote sensing4675 of 75
Vector and overlay4618 of 18
Networks and routing535 of 5
Climate and time series5624 of 24
Spatial statistics and interpolation2924 of 24
Mapping and cartography5065 of 65
Recognising infeasible requests4879 of 79

Terminus-2 – Gemini 3.1 Flash-Lite (low): score by data type

Average pass@1 over the tasks of each data type, 0 to 100 · Higher is better

Terminus-2 – Gemini 3.1 Flash-Lite (low): score by data typeOptical multispectral: No data; Elevation or terrain: No data; Climate or gridded time series: No data; Other gridded data: No data; Vector only: No data; Tabular only: No data0255075100No dataOpticalmultispectralNo dataElevation orterrainNo dataClimate orgridded timeseriesNo dataOther griddeddataNo dataVector onlyNo dataTabular only
How to read this chart

What this shows. The same attempts as the index, grouped by data type. Groups with few tasks are less certain. About these breakdowns.

Data table: Terminus-2 – Gemini 3.1 Flash-Lite (low): score by data type
Terminus-2 – Gemini 3.1 Flash-Lite (low): by data type
Data typeScoreTasks with a decided attempt
Optical multispectralNo data0 of 43
Elevation or terrainNo data0 of 16
Climate or gridded time seriesNo data0 of 17
Other gridded dataNo data0 of 26
Vector onlyNo data0 of 108
Tabular onlyNo data0 of 2

Terminus-2 – Gemini 3.1 Flash-Lite (low): score by task length

Average pass@1 over the tasks of each task length, 0 to 100 · Higher is better

Terminus-2 – Gemini 3.1 Flash-Lite (low): score by task length0255075100
How to read this chart

What this shows. The same attempts as the index, grouped by task length. Groups with few tasks are less certain. About these breakdowns.

Data table: Terminus-2 – Gemini 3.1 Flash-Lite (low): score by task length
Terminus-2 – Gemini 3.1 Flash-Lite (low): by task length
Task lengthScoreTasks with a decided attempt

Compared with other models

Geospatial Agent Index

Equal-weight average of the evaluation scores, 0 to 100 · Higher is better

  • Partial coverage
Geospatial Agent IndexClaude Code – Claude Opus 5.5 (xhigh): 82; Claude Code – Claude Opus 5.5 (high): 82* (partial coverage); Claude Code – Claude Sonnet 5.5 (high): 81* (partial coverage); Claude Code – Claude Opus 5.5 (medium): 80* (partial coverage); Codex – GPT-6.1 Sol (high): 80; Codex – GPT-6.1 Sol (medium): 79* (partial coverage); Codex – GPT-6.1 Sol (xhigh): 78* (partial coverage); Claude Code – Claude Haiku 5.5 (xhigh): 78; Codex – GPT-6 Luna (xhigh): 77* (partial coverage); Claude Code – Claude Haiku 5.5 (medium): 77* (partial coverage); Claude Code – Claude Haiku 5.5 (high): 77* (partial coverage); Claude Code – Claude Sonnet 5.5 (medium): 76* (partial coverage); Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high): 75* (partial coverage); Terminus-2 – Gemini 3.1 Pro (preview) (high): 75; Terminus-2 – DeepSeek V4 Flash (max): 74* (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 74; Claude Code – Claude Sonnet 5.5 (xhigh): 74; Terminus-2 – GLM-5.2 (max): 74* (partial coverage); Terminus-2 – Kimi K2.7 Code (always on): 74; Terminus-2 – GLM-5.2 (high): 72* (partial coverage); Terminus-2 – Kimi K2.6 (always on): 72; Terminus-2 – GLM-5.3 (low): 72* (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 71; Terminus-2 – GLM-5.2 (none): 70* (partial coverage); Terminus-2 – GLM-5.3 (max, default): 68* (partial coverage); Terminus-2 – GLM-5.3-Flash (low): 66* (partial coverage); Terminus-2 – Gemini 3.1 Pro (preview) (low): 65; Terminus-2 – Gemma 4 26B A4B (thinking on): 62; Terminus-2 – Gemini 3.1 Flash-Lite (high): 59* (partial coverage); Terminus-2 – GLM-5.3-Flash (max): 56* (partial coverage); Terminus-2 – gpt-oss-120b (high): 52; Terminus-2 – Gemini 3.1 Flash-Lite (low): 49; Terminus-2 – GLM-4.7-Flash (thinking on): 40; Terminus-2 – gpt-oss-120b (low): 37; Terminus-2 – Gemini 3.1 Flash-Lite (minimal): 34* (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 62* (partial coverage); Terminus-2 – Mistral Small 3.1 24B (Non-reasoning): 0* (partial coverage)025507510082Claude CodeClaude Opus5.5 (xhigh)82*Claude CodeClaude Opus5.5 (high)81*Claude CodeClaude Sonnet5.5 (high)80*Claude CodeClaude Opus5.5 (medium)80CodexGPT-6.1 Sol(high)79*CodexGPT-6.1 Sol(medium)78*CodexGPT-6.1 Sol(xhigh)78Claude CodeClaude Haiku5.5 (xhigh)77*CodexGPT-6 Luna(xhigh)77*Claude CodeClaude Haiku5.5 (medium)77*Claude CodeClaude Haiku5.5 (high)76*Claude CodeClaude Sonnet5.5 (medium)75*Terminus-2Claude Sonnet4.6(adaptive,75Terminus-2Gemini 3.1Pro (preview)(high)74*Terminus-2DeepSeek V4Flash (max)74Terminus-2DeepSeek V4Pro (high)74Claude CodeClaude Sonnet5.5 (xhigh)74*Terminus-2GLM-5.2 (max)74Terminus-2Kimi K2.7Code (alwayson)72*Terminus-2GLM-5.2(high)72Terminus-2Kimi K2.6(always on)72*Terminus-2GLM-5.3 (low)71Terminus-2DeepSeek V4Flash (high)70*Terminus-2GLM-5.2(none)68*Terminus-2GLM-5.3 (max,default)66*Terminus-2GLM-5.3-Flash(low)65Terminus-2Gemini 3.1Pro (preview)(low)62Terminus-2Gemma 4 26BA4B (thinkingon)59*Terminus-2Gemini 3.1Flash-Lite(high)56*Terminus-2GLM-5.3-Flash(max)52Terminus-2gpt-oss-120b(high)49Terminus-2Gemini 3.1Flash-Lite(low)40Terminus-2GLM-4.7-Flash(thinking on)37Terminus-2gpt-oss-120b(low)34*Terminus-2Gemini 3.1Flash-Lite(minimal)62*Terminus-2Qwen 3.8 27B(xhigh)0*Terminus-2Mistral Small3.1 24B(Non-reasoning)
How to read this chart

What this metric means. The equal-weight average of the 4 evaluation scores. Each evaluation score is the average pass@1 over that benchmark's tasks, from 0 to 100.

Not yet ranked. Entries without a score on every benchmark come after the ranked entries and are not ranked: their index averages only the benchmarks they have.

Partial coverage. A hatched bar, and an asterisk in tables, marks an entry that does not yet have results for every planned attempt. While attempts await review, the table gives the range the score could still reach.

Attempts per task. Each task's score is the average of its recorded attempts (up to 3). Some entries and tracks currently have 1 attempt per task; further attempts will be added and averaged in. Each entry shows its attempts per task.

Data table: Geospatial Agent Index
Geospatial Agent Index
ModelCreatorGeospatial Agent IndexRange while attempts are pendingBenchmarks coveredCoverage (attempts)
Claude Code – Claude Opus 5.5 (xhigh)Anthropic82 (attempts per task: 1)No pending judgements4 of 4290 of 290 planned attempts
Claude Code – Claude Opus 5.5 (high)Anthropic82* (attempts per task: 2.5)63 to 844 of 4531 of 870 planned attempts
Claude Code – Claude Sonnet 5.5 (high)Anthropic81* (attempts per task: 2)77 to 814 of 4556 of 870 planned attempts
Claude Code – Claude Opus 5.5 (medium)Anthropic80* (attempts per task: 2.7)58 to 834 of 4573 of 870 planned attempts
Codex – GPT-6.1 Sol (high)OpenAI80 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Codex – GPT-6.1 Sol (medium)OpenAI79* (attempts per task: 2.6)No pending judgements; coverage incomplete4 of 4749 of 870 planned attempts (6 excluded)
Codex – GPT-6.1 Sol (xhigh)OpenAI78* (attempts per task: 2.6)69 to 784 of 4668 of 870 planned attempts
Claude Code – Claude Haiku 5.5 (xhigh)Anthropic78 (attempts per task: 1)No pending judgements4 of 4290 of 290 planned attempts
Codex – GPT-6 Luna (xhigh)OpenAI77* (attempts per task: 2.4)69 to 784 of 4627 of 870 planned attempts
Claude Code – Claude Haiku 5.5 (medium)Anthropic77* (attempts per task: 2)74 to 784 of 4538 of 870 planned attempts
Claude Code – Claude Haiku 5.5 (high)Anthropic77* (attempts per task: 2.3)61 to 794 of 4526 of 870 planned attempts
Claude Code – Claude Sonnet 5.5 (medium)Anthropic76* (attempts per task: 2.2)71 to 764 of 4580 of 870 planned attempts
Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high)Anthropic75* (attempts per task: 1.2)74 to 754 of 4335 of 870 planned attempts (138 excluded)
Terminus-2 – Gemini 3.1 Pro (preview) (high)Google75 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – DeepSeek V4 Flash (max)DeepSeek74* (attempts per task: 1.2)73 to 754 of 4344 of 870 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek74 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Claude Code – Claude Sonnet 5.5 (xhigh)Anthropic74 (attempts per task: 1)No pending judgements4 of 4290 of 290 planned attempts
Terminus-2 – GLM-5.2 (max)Z.ai74* (attempts per task: 1.2)74 to 744 of 4338 of 870 planned attempts (1 excluded)
Terminus-2 – Kimi K2.7 Code (always on)Moonshot AI74 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – GLM-5.2 (high)Z.ai72* (attempts per task: 1.1)71 to 724 of 4310 of 870 planned attempts (1 excluded)
Terminus-2 – Kimi K2.6 (always on)Moonshot AI72 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – GLM-5.3 (low)Z.ai72* (attempts per task: 1)70 to 724 of 4285 of 290 planned attempts
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek71 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – GLM-5.2 (none)Z.ai70* (attempts per task: 1.6)68 to 704 of 4446 of 870 planned attempts
Terminus-2 – GLM-5.3 (max, default)Z.ai68* (attempts per task: 1.2)67 to 684 of 4335 of 870 planned attempts
Terminus-2 – GLM-5.3-Flash (low)Z.ai66* (attempts per task: 1.4)61 to 674 of 4380 of 870 planned attempts (4 excluded)
Terminus-2 – Gemini 3.1 Pro (preview) (low)Google65 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts (22 excluded)
Terminus-2 – Gemma 4 26B A4B (thinking on)Google62 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – Gemini 3.1 Flash-Lite (high)Google59* (attempts per task: 3)59 to 594 of 4864 of 870 planned attempts (6 excluded)
Terminus-2 – GLM-5.3-Flash (max)Z.ai56* (attempts per task: 1.2)55 to 564 of 4342 of 870 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI52 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – Gemini 3.1 Flash-Lite (low)Google49 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – GLM-4.7-Flash (thinking on)Z.ai40 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – gpt-oss-120b (low)OpenAI37 (attempts per task: 3)No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – Gemini 3.1 Flash-Lite (minimal)Google34* (attempts per task: 3)No pending judgements; coverage incomplete4 of 4869 of 870 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)62* (attempts per task: 1 on 101 of 290 tasks so far)No pending judgements; coverage incomplete3 of 4101 of 870 planned attempts
Terminus-2 – Mistral Small 3.1 24B (Non-reasoning)Mistral AI0* (attempts per task: 1 on 9 of 290 tasks so far)No pending judgements; coverage incomplete1 of 49 of 351 planned attempts

Geospatial Agent Index vs. cost per task

Higher and further left is better

  • Partial coverage
  • Most attractive quadrant
  • Pareto line
Geospatial Agent Index vs. cost per taskTerminus-2 – Claude Sonnet 4.6 (adaptive, effort high): Geospatial Agent Index 74.82, cost per task (usd) $0.272; Terminus-2 – DeepSeek V4 Flash (max): Geospatial Agent Index 74.5, cost per task (usd) $0.067; Terminus-2 – DeepSeek V4 Flash (high): Geospatial Agent Index 70.76, cost per task (usd) $0.067; Terminus-2 – DeepSeek V4 Pro (high): Geospatial Agent Index 74.2, cost per task (usd) $0.273; Terminus-2 – Gemini 3.1 Flash-Lite (low): Geospatial Agent Index 49.45, cost per task (usd) $0.012; Terminus-2 – Gemini 3.1 Flash-Lite (minimal): Geospatial Agent Index 34.37, cost per task (usd) $0.0056; Terminus-2 – Gemini 3.1 Flash-Lite (high): Geospatial Agent Index 58.87, cost per task (usd) $0.021; Terminus-2 – Gemini 3.1 Pro (preview) (low): Geospatial Agent Index 65.17, cost per task (usd) $0.061; Terminus-2 – Gemini 3.1 Pro (preview) (high): Geospatial Agent Index 74.77, cost per task (usd) $0.252; Terminus-2 – Gemma 4 26B A4B (thinking on): Geospatial Agent Index 61.62, cost per task (usd) $0.016; Terminus-2 – GLM-4.7-Flash (thinking on): Geospatial Agent Index 39.77, cost per task (usd) $0.030; Terminus-2 – GLM-5.2 (high): Geospatial Agent Index 72.42, cost per task (usd) $0.231; Terminus-2 – GLM-5.2 (none): Geospatial Agent Index 69.92, cost per task (usd) $0.102; Terminus-2 – GLM-5.2 (max): Geospatial Agent Index 74, cost per task (usd) $0.359; Terminus-2 – GLM-5.3-Flash (low): Geospatial Agent Index 65.66, cost per task (usd) $0.028; Terminus-2 – GLM-5.3-Flash (max): Geospatial Agent Index 55.86, cost per task (usd) $0.042; Terminus-2 – GLM-5.3 (low): Geospatial Agent Index 71.75, cost per task (usd) $0.251; Terminus-2 – GLM-5.3 (max, default): Geospatial Agent Index 68.29, cost per task (usd) $0.360; Terminus-2 – gpt-oss-120b (low): Geospatial Agent Index 37.05, cost per task (usd) $0.0088; Terminus-2 – gpt-oss-120b (high): Geospatial Agent Index 52.1, cost per task (usd) $0.026; Terminus-2 – Kimi K2.6 (always on): Geospatial Agent Index 72.4, cost per task (usd) $0.072; Terminus-2 – Kimi K2.7 Code (always on): Geospatial Agent Index 73.7, cost per task (usd) $0.119; Terminus-2 – Mistral Small 3.1 24B (Non-reasoning): Geospatial Agent Index 0, cost per task (usd) $0.247; Terminus-2 – Qwen 3.8 27B (xhigh): Geospatial Agent Index 62.5, cost per task (usd) $0.066; Claude Code – Claude Opus 5.5 (medium): Geospatial Agent Index 80.36, cost per task (usd) $0.137; Claude Code – Claude Opus 5.5 (high): Geospatial Agent Index 82.3, cost per task (usd) $0.167; Claude Code – Claude Sonnet 5.5 (medium): Geospatial Agent Index 75.61, cost per task (usd) $0.053; Claude Code – Claude Sonnet 5.5 (high): Geospatial Agent Index 80.69, cost per task (usd) $0.062; Claude Code – Claude Haiku 5.5 (medium): Geospatial Agent Index 77.17, cost per task (usd) $0.0074; Claude Code – Claude Haiku 5.5 (high): Geospatial Agent Index 76.52, cost per task (usd) $0.010; Claude Code – Claude Opus 5.5 (xhigh): Geospatial Agent Index 82.32, cost per task (usd) $0.375; Claude Code – Claude Sonnet 5.5 (xhigh): Geospatial Agent Index 74.18, cost per task (usd) $0.163; Claude Code – Claude Haiku 5.5 (xhigh): Geospatial Agent Index 77.89, cost per task (usd) $0.023; Codex – GPT-6.1 Sol (medium): Geospatial Agent Index 79.48, cost per task (usd) $0.060; Codex – GPT-6.1 Sol (high): Geospatial Agent Index 80.16, cost per task (usd) $0.080; Codex – GPT-6.1 Sol (xhigh): Geospatial Agent Index 77.94, cost per task (usd) $0.121; Codex – GPT-6 Luna (xhigh): Geospatial Agent Index 77.21, cost per task (usd) $0.0092Most attractive quadrant0255075100$0.0010$0.0020$0.0050$0.010$0.020$0.050$0.100$0.200$0.500$1.00Cost per task (USD) (log scale)Geospatial Agent IndexClaude Code – Claude Opus 5.5 (xhigh)Claude Code – Claude Opus 5.5 (high)Claude Code – Claude Sonnet 5.5 (high)Claude Code – Claude Opus 5.5 (medium)Codex – GPT-6.1 Sol (high)Codex – GPT-6.1 Sol (medium)Codex – GPT-6.1 Sol (xhigh)Claude Code – Claude Haiku 5.5 (xhigh)Codex – GPT-6 Luna (xhigh)Claude Code – Claude Haiku 5.5 (medium)Claude Code – Claude Haiku 5.5 (high)Claude Code – Claude Sonnet 5.5 (medium)Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high)Terminus-2 – Gemini 3.1 Pro (preview) (high)Terminus-2 – DeepSeek V4 Flash (max)Terminus-2 – DeepSeek V4 Pro (high)Claude Code – Claude Sonnet 5.5 (xhigh)Terminus-2 – GLM-5.2 (max)Terminus-2 – Kimi K2.7 Code (always on)Terminus-2 – GLM-5.2 (high)Terminus-2 – Kimi K2.6 (always on)Terminus-2 – GLM-5.3 (low)Terminus-2 – DeepSeek V4 Flash (high)Terminus-2 – GLM-5.2 (none)Terminus-2 – GLM-5.3 (max, default)Terminus-2 – GLM-5.3-Flash (low)Terminus-2 – Gemini 3.1 Pro (preview) (low)Terminus-2 – Qwen 3.8 27B (xhigh)Terminus-2 – Gemma 4 26B A4B (thinking on)Terminus-2 – Gemini 3.1 Flash-Lite (high)Terminus-2 – GLM-5.3-Flash (max)Terminus-2 – gpt-oss-120b (high)Terminus-2 – Gemini 3.1 Flash-Lite (low)Terminus-2 – GLM-4.7-Flash (thinking on)Terminus-2 – gpt-oss-120b (low)Terminus-2 – Gemini 3.1 Flash-Lite (minimal)Terminus-2 – Mistral Small 3.1 24B (Non-reasoning)
Data table: Geospatial Agent Index vs. cost per task
Geospatial Agent Index vs. cost per task
ModelGeospatial Agent IndexCost per task (USD)On the Pareto line
Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high)75* (attempts per task: 1.2)$0.272No
Terminus-2 – DeepSeek V4 Flash (max)74* (attempts per task: 1.2)$0.067No
Terminus-2 – DeepSeek V4 Flash (high)71 (attempts per task: 3)$0.067No
Terminus-2 – DeepSeek V4 Pro (high)74 (attempts per task: 3)$0.273No
Terminus-2 – Gemini 3.1 Flash-Lite (low)49 (attempts per task: 3)$0.012No
Terminus-2 – Gemini 3.1 Flash-Lite (minimal)34* (attempts per task: 3)$0.0056Yes
Terminus-2 – Gemini 3.1 Flash-Lite (high)59* (attempts per task: 3)$0.021No
Terminus-2 – Gemini 3.1 Pro (preview) (low)65 (attempts per task: 3)$0.061No
Terminus-2 – Gemini 3.1 Pro (preview) (high)75 (attempts per task: 3)$0.252No
Terminus-2 – Gemma 4 26B A4B (thinking on)62 (attempts per task: 3)$0.016No
Terminus-2 – GLM-4.7-Flash (thinking on)40 (attempts per task: 3)$0.030No
Terminus-2 – GLM-5.2 (high)72* (attempts per task: 1.1)$0.231No
Terminus-2 – GLM-5.2 (none)70* (attempts per task: 1.6)$0.102No
Terminus-2 – GLM-5.2 (max)74* (attempts per task: 1.2)$0.359No
Terminus-2 – GLM-5.3-Flash (low)66* (attempts per task: 1.4)$0.028No
Terminus-2 – GLM-5.3-Flash (max)56* (attempts per task: 1.2)$0.042No
Terminus-2 – GLM-5.3 (low)72* (attempts per task: 1)$0.251No
Terminus-2 – GLM-5.3 (max, default)68* (attempts per task: 1.2)$0.360No
Terminus-2 – gpt-oss-120b (low)37 (attempts per task: 3)$0.0088No
Terminus-2 – gpt-oss-120b (high)52 (attempts per task: 3)$0.026No
Terminus-2 – Kimi K2.6 (always on)72 (attempts per task: 3)$0.072No
Terminus-2 – Kimi K2.7 Code (always on)74 (attempts per task: 3)$0.119No
Terminus-2 – Mistral Small 3.1 24B (Non-reasoning)0* (attempts per task: 1 on 9 of 290 tasks so far)$0.247No
Terminus-2 – Qwen 3.8 27B (xhigh)62* (attempts per task: 1 on 101 of 290 tasks so far)$0.066No
Claude Code – Claude Opus 5.5 (medium)80* (attempts per task: 2.7)$0.137No
Claude Code – Claude Opus 5.5 (high)82* (attempts per task: 2.5)$0.167Yes
Claude Code – Claude Sonnet 5.5 (medium)76* (attempts per task: 2.2)$0.053No
Claude Code – Claude Sonnet 5.5 (high)81* (attempts per task: 2)$0.062Yes
Claude Code – Claude Haiku 5.5 (medium)77* (attempts per task: 2)$0.0074Yes
Claude Code – Claude Haiku 5.5 (high)77* (attempts per task: 2.3)$0.010No
Claude Code – Claude Opus 5.5 (xhigh)82 (attempts per task: 1)$0.375Yes
Claude Code – Claude Sonnet 5.5 (xhigh)74 (attempts per task: 1)$0.163No
Claude Code – Claude Haiku 5.5 (xhigh)78 (attempts per task: 1)$0.023Yes
Codex – GPT-6.1 Sol (medium)79* (attempts per task: 2.6)$0.060Yes
Codex – GPT-6.1 Sol (high)80 (attempts per task: 3)$0.080No
Codex – GPT-6.1 Sol (xhigh)78* (attempts per task: 2.6)$0.121No
Codex – GPT-6 Luna (xhigh)77* (attempts per task: 2.4)$0.0092Yes

Head-to-head comparisons

Specification and settings

Developer
Google
Context window
1,048,576 tokens
Image input
Yes
Reasoning setting
low
Temperature
1.0
Maximum output tokens
65,536
Input price per 1M tokens
$0.25
Cached input price per 1M tokens
$0.025
Output price per 1M tokens
$1.50

Why this model was chosen: model selection.