Terminus-2 – Gemini 3.8 Flash (high): geospatial agent results
Terminus-2 – Gemini 3.8 Flash (high) by Google, run under Terminus-2: its Geospatial Agent Index, score on each benchmark and each kind of geospatial work, cost, time and token use.
Summary
- Geospatial Agent Index
- 39* (attempts per task: 2.9) (rank 34 of 37)
- Cost per task
- $0.308
- Time per task
- 8.2 min
- Tokens per attempt
- 982k
- Turns per attempt
- 32.4
- Attempts decided
- 278
Score by benchmark
Terminus-2 – Gemini 3.8 Flash (high): score by benchmark
Average pass@1, 0 to 100 · Higher is better
Data table: Terminus-2 – Gemini 3.8 Flash (high): score by benchmark
| Benchmark | Score | Tasks with a decided attempt | Cost per task | Time per task |
|---|---|---|---|---|
| GeoAgentBench | 58 | 50 of 50 | $0.187 | 6.2 min |
| GeoBenchX | 40 | 173 of 173 | $0.340 | 8.8 min |
| Earth-Bench | 28 | 36 of 48 | $0.341 | 8.6 min |
| GeoAnalystBench | 32 | 19 of 19 | $0.274 | 7.6 min |
Score by task format and kind of work
Terminus-2 – Gemini 3.8 Flash (high): score by task format
Average pass@1 over the tasks of each task format, 0 to 100 · Higher is better
How to read this chart
What this shows. The same attempts as the index, grouped by task format. Groups with few tasks are less certain. About these breakdowns.
Data table: Terminus-2 – Gemini 3.8 Flash (high): score by task format
| Task format | Score | Tasks with a decided attempt |
|---|---|---|
| Multi-step analysis | 33 | 151 of 151 |
| Single-answer questions | 25 | 48 of 159 |
| Recognising infeasible requests | 67 | 79 of 79 |
Terminus-2 – Gemini 3.8 Flash (high): score by kind of work
Average pass@1 over the tasks of each kind of work, 0 to 100 · Higher is better
How to read this chart
What this shows. The same attempts as the index, grouped by kind of work. Groups with few tasks are less certain. About these breakdowns.
Data table: Terminus-2 – Gemini 3.8 Flash (high): score by kind of work
| Kind of work | Score | Tasks with a decided attempt |
|---|---|---|
| Raster and remote sensing | 37 | 68 of 123 |
| Vector and overlay | 44 | 18 of 18 |
| Networks and routing | 80 | 5 of 5 |
| Climate and time series | 26 | 19 of 75 |
| Spatial statistics and interpolation | 17 | 24 of 24 |
| Mapping and cartography | 25 | 65 of 65 |
| Recognising infeasible requests | 67 | 79 of 79 |
Terminus-2 – Gemini 3.8 Flash (high): score by data type
Average pass@1 over the tasks of each data type, 0 to 100 · Higher is better
How to read this chart
What this shows. The same attempts as the index, grouped by data type. Groups with few tasks are less certain. About these breakdowns.
Data table: Terminus-2 – Gemini 3.8 Flash (high): score by data type
| Data type | Score | Tasks with a decided attempt |
|---|---|---|
| Optical multispectral | 18 | 33 of 117 |
| Elevation or terrain | 56 | 16 of 16 |
| Climate or gridded time series | 24 | 17 of 18 |
| Other gridded data | 52 | 23 of 61 |
| Vector only | 29 | 108 of 108 |
| Tabular only | 50 | 2 of 2 |
Terminus-2 – Gemini 3.8 Flash (high): score by task length
Average pass@1 over the tasks of each task length, 0 to 100 · Higher is better
How to read this chart
What this shows. The same attempts as the index, grouped by task length. Groups with few tasks are less certain. About these breakdowns.
Data table: Terminus-2 – Gemini 3.8 Flash (high): score by task length
| Task length | Score | Tasks with a decided attempt |
|---|---|---|
| Short tasks | 75 | 87 of 129 |
| Medium tasks | 36 | 107 of 131 |
| Long tasks | 14 | 84 of 129 |
Compared with other models
Geospatial Agent Index
Equal-weight average of the evaluation scores, 0 to 100 · Higher is better
- Partial coverage
How to read this chart
What this metric means. The equal-weight average of the 4 evaluation scores. Each evaluation score is the average pass@1 over that benchmark's tasks, from 0 to 100.
Not yet ranked. Entries without a score on every benchmark come after the ranked entries and are not ranked: their index averages only the benchmarks they have.
Partial coverage. A hatched bar, and an asterisk in tables, marks an entry that does not yet have results for every planned attempt. While attempts await review, the table gives the range the score could still reach.
Attempts per task. Each task's score is the average of its recorded attempts (up to 3). Some entries and tracks currently have 1 attempt per task; further attempts will be added and averaged in. Each entry shows its attempts per task.
Data table: Geospatial Agent Index
| Model | Creator | Geospatial Agent Index | Range while attempts are pending | Benchmarks covered | Coverage (attempts) |
|---|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 72* (attempts per task: 3.8) | 30 to 84 | 4 of 4 | 544 of 1,167 planned attempts |
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 72* (attempts per task: 3.2) | 30 to 84 | 4 of 4 | 368 of 1,167 planned attempts (21 excluded) |
| Terminus-2 – GLM-5.2 (high) | Z.ai | 72* (attempts per task: 1.8) | 68 to 73 | 4 of 4 | 484 of 1,167 planned attempts (3 excluded) |
| Terminus-2 – DeepSeek V4 Flash (max) | DeepSeek | 71* (attempts per task: 1.9) | 65 to 73 | 4 of 4 | 464 of 1,167 planned attempts (19 excluded) |
| Codex – GPT-6.1 Sol (medium) | OpenAI | 70* (attempts per task: 3.3) | 67 to 80 | 4 of 4 | 870 of 1,167 planned attempts (1 excluded) |
| Codex – GPT-6.1 Sol (high) | OpenAI | 69* (attempts per task: 3.3) | 66 to 78 | 4 of 4 | 877 of 1,167 planned attempts |
| Terminus-2 – GLM-5.2 (max) | Z.ai | 69* (attempts per task: 2) | 65 to 71 | 4 of 4 | 508 of 1,167 planned attempts (3 excluded) |
| Claude Code – Claude Opus 5.5 (high) | Anthropic | 69* (attempts per task: 3.9) | 66 to 80 | 4 of 4 | 921 of 1,167 planned attempts |
| Codex – GPT-6 Luna (xhigh) | OpenAI | 69* (attempts per task: 3.3) | 34 to 86 | 4 of 4 | 444 of 1,167 planned attempts |
| Codex – GPT-6.1 Sol (xhigh) | OpenAI | 68* (attempts per task: 3.3) | 66 to 78 | 4 of 4 | 871 of 1,167 planned attempts |
| Claude Code – Claude Opus 5.5 (xhigh) | Anthropic | 68* (attempts per task: 1.3) | 65 to 78 | 4 of 4 | 308 of 683 planned attempts |
| Claude Code – Claude Haiku 5.5 (xhigh) | Anthropic | 68* (attempts per task: 1.3) | 66 to 78 | 4 of 4 | 319 of 683 planned attempts |
| Terminus-2 – Kimi K2.7 Code (always on) | Moonshot AI | 68* (attempts per task: 3.7) | 28 to 83 | 4 of 4 | 496 of 1,167 planned attempts |
| Claude Code – Claude Sonnet 5.5 (high) | Anthropic | 68* (attempts per task: 4) | 66 to 79 | 4 of 4 | 949 of 1,167 planned attempts |
| Claude Code – Claude Opus 5.5 (medium) | Anthropic | 67* (attempts per task: 4) | 66 to 78 | 4 of 4 | 942 of 1,167 planned attempts |
| Terminus-2 – Kimi K2.6 (always on) | Moonshot AI | 67* (attempts per task: 3.8) | 27 to 82 | 4 of 4 | 532 of 1,167 planned attempts |
| Claude Code – Claude Haiku 5.5 (high) | Anthropic | 66* (attempts per task: 3.8) | 64 to 76 | 4 of 4 | 937 of 1,167 planned attempts |
| Claude Code – Claude Haiku 5.5 (medium) | Anthropic | 65* (attempts per task: 4) | 47 to 80 | 4 of 4 | 745 of 1,167 planned attempts |
| Claude Code – Claude Sonnet 5.5 (medium) | Anthropic | 65* (attempts per task: 4) | 65 to 77 | 4 of 4 | 951 of 1,167 planned attempts |
| Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high) | Anthropic | 65* (attempts per task: 1.1) | 62 to 66 | 4 of 4 | 295 of 1,167 planned attempts (149 excluded) |
| Terminus-2 – Gemini 3.1 Pro (preview) (high) | 65* (attempts per task: 2.9) | No pending judgements; coverage incomplete | 4 of 4 | 834 of 1,167 planned attempts | |
| Terminus-2 – Gemini 3.8 Flash (low) | 64* (attempts per task: 2.9) | No pending judgements; coverage incomplete | 4 of 4 | 833 of 1,167 planned attempts (74 excluded) | |
| Claude Code – Claude Sonnet 5.5 (xhigh) | Anthropic | 64* (attempts per task: 1.3) | 62 to 74 | 4 of 4 | 318 of 683 planned attempts |
| Terminus-2 – GLM-5.3 (low) | Z.ai | 64* (attempts per task: 1.8) | 63 to 64 | 4 of 4 | 513 of 1,167 planned attempts |
| Terminus-2 – GLM-5.3 (max, default) | Z.ai | 64* (attempts per task: 1.8) | 60 to 65 | 4 of 4 | 463 of 1,167 planned attempts |
| Terminus-2 – GLM-5.2 (none) | Z.ai | 64* (attempts per task: 2.9) | 50 to 69 | 4 of 4 | 636 of 1,167 planned attempts (2 excluded) |
| Terminus-2 – Gemini 3.1 Pro (preview) (low) | 63* (attempts per task: 2.9) | 21 to 88 | 4 of 4 | 278 of 1,167 planned attempts (21 excluded) | |
| Terminus-2 – GLM-5.3-Flash (low) | Z.ai | 62* (attempts per task: 1.6) | 45 to 72 | 4 of 4 | 298 of 1,167 planned attempts (10 excluded) |
| Terminus-2 – GLM-5.3-Flash (max) | Z.ai | 56* (attempts per task: 1.9) | 51 to 59 | 4 of 4 | 493 of 1,167 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (thinking on) | 56* (attempts per task: 3.7) | 22 to 77 | 4 of 4 | 516 of 1,167 planned attempts | |
| Terminus-2 – Gemini 3.1 Flash-Lite (low) | 51* (attempts per task: 2.9) | 17 to 84 | 4 of 4 | 278 of 1,167 planned attempts | |
| Terminus-2 – Gemini 3.1 Flash-Lite (high) | 51* (attempts per task: 2.9) | 17 to 84 | 4 of 4 | 272 of 1,167 planned attempts (6 excluded) | |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 48* (attempts per task: 4) | 20 to 74 | 4 of 4 | 611 of 1,167 planned attempts |
| Terminus-2 – Gemini 3.8 Flash (high) | 39* (attempts per task: 2.9) | 13 to 80 | 4 of 4 | 278 of 1,167 planned attempts (1 excluded) | |
| Terminus-2 – gpt-oss-120b (low) | OpenAI | 38* (attempts per task: 4) | 15 to 69 | 4 of 4 | 611 of 1,167 planned attempts |
| Terminus-2 – Gemini 3.1 Flash-Lite (minimal) | 33* (attempts per task: 2.9) | 11 to 78 | 4 of 4 | 278 of 1,167 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (thinking on) | Z.ai | 33* (attempts per task: 3.6) | 14 to 69 | 4 of 4 | 483 of 1,167 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 55* (attempts per task: 1 on 89 of 290 tasks so far) | No pending judgements; coverage incomplete | 3 of 4 | 89 of 1,167 planned attempts |
| Terminus-2 – Mistral Small 3.1 24B (Non-reasoning) | Mistral AI | 0* (attempts per task: 1 on 9 of 290 tasks so far) | No pending judgements; coverage incomplete | 1 of 4 | 9 of 648 planned attempts (4 excluded) |
Geospatial Agent Index vs. cost per task
Higher and further left is better
- Partial coverage
- Most attractive quadrant
- Pareto line
How to read this chart
Not yet ranked. Not plotted, so the axes fit the ranked entries; listed in the table: Terminus-2 – Mistral Small 3.1 24B (Non-reasoning), Terminus-2 – Qwen 3.8 27B (xhigh).
What cost is measuring. What one attempt at a task would cost if paid token by token at the provider's published API prices, averaged over attempts: ordinary input tokens, input tokens served from the provider's cache at the lower cached price, tokens written to the cache where the provider charges for that, and output tokens (thinking included). Cost assumes each harness uses the model's pay-per-token API, priced at public list rates; sandboxes and our own engineering time are not included. Failed attempts count too.
How to read this chart. Each point is one model configuration. Further left means a lower average cost per task (log scale); higher means a higher Geospatial Agent Index. Both axes fit the models shown, so the score axis does not start at zero. The shaded most attractive quadrant is the top-left quarter of the plot, split at the middle of each axis: a reading aid that moves with the axes when you change the models shown, not a fixed bar. The dotted Pareto line joins the shown models that no cheaper shown model beats on score.
Data table: Geospatial Agent Index vs. cost per task
| Model | Geospatial Agent Index | Cost per task (USD) | On the Pareto line |
|---|---|---|---|
| Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high) | 65* (attempts per task: 1.1) | $0.252 | No |
| Terminus-2 – DeepSeek V4 Flash (max) | 71* (attempts per task: 1.9) | $0.067 | Yes |
| Terminus-2 – DeepSeek V4 Flash (high) | 72* (attempts per task: 3.2) | $0.070 | Yes |
| Terminus-2 – DeepSeek V4 Pro (high) | 72* (attempts per task: 3.8) | $0.365 | Yes |
| Terminus-2 – Gemini 3.1 Flash-Lite (low) | 51* (attempts per task: 2.9) | $0.011 | No |
| Terminus-2 – Gemini 3.1 Flash-Lite (minimal) | 33* (attempts per task: 2.9) | $0.0054 | Yes |
| Terminus-2 – Gemini 3.1 Flash-Lite (high) | 51* (attempts per task: 2.9) | $0.020 | No |
| Terminus-2 – Gemini 3.1 Pro (preview) (low) | 63* (attempts per task: 2.9) | $0.060 | No |
| Terminus-2 – Gemini 3.1 Pro (preview) (high) | 65* (attempts per task: 2.9) | $0.237 | No |
| Terminus-2 – Gemini 3.8 Flash (high) | 39* (attempts per task: 2.9) | $0.308 | No |
| Terminus-2 – Gemini 3.8 Flash (low) | 64* (attempts per task: 2.9) | $0.115 | No |
| Terminus-2 – Gemma 4 26B A4B (thinking on) | 56* (attempts per task: 3.7) | $0.017 | No |
| Terminus-2 – GLM-4.7-Flash (thinking on) | 33* (attempts per task: 3.6) | $0.033 | No |
| Terminus-2 – GLM-5.2 (high) | 72* (attempts per task: 1.8) | $0.255 | No |
| Terminus-2 – GLM-5.2 (none) | 64* (attempts per task: 2.9) | $0.104 | No |
| Terminus-2 – GLM-5.2 (max) | 69* (attempts per task: 2) | $0.361 | No |
| Terminus-2 – GLM-5.3-Flash (low) | 62* (attempts per task: 1.6) | $0.027 | No |
| Terminus-2 – GLM-5.3-Flash (max) | 56* (attempts per task: 1.9) | $0.050 | No |
| Terminus-2 – GLM-5.3 (low) | 64* (attempts per task: 1.8) | $0.326 | No |
| Terminus-2 – GLM-5.3 (max, default) | 64* (attempts per task: 1.8) | $0.434 | No |
| Terminus-2 – gpt-oss-120b (low) | 38* (attempts per task: 4) | $0.0080 | No |
| Terminus-2 – gpt-oss-120b (high) | 48* (attempts per task: 4) | $0.031 | No |
| Terminus-2 – Kimi K2.6 (always on) | 67* (attempts per task: 3.8) | $0.093 | No |
| Terminus-2 – Kimi K2.7 Code (always on) | 68* (attempts per task: 3.7) | $0.142 | No |
| Terminus-2 – Mistral Small 3.1 24B (Non-reasoning) | 0* (attempts per task: 1 on 9 of 290 tasks so far) | $0.215 | Not plotted: not yet ranked |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 55* (attempts per task: 1 on 89 of 290 tasks so far) | $0.063 | Not plotted: not yet ranked |
| Claude Code – Claude Opus 5.5 (medium) | 67* (attempts per task: 4) | $0.139 | No |
| Claude Code – Claude Opus 5.5 (high) | 69* (attempts per task: 3.9) | $0.177 | No |
| Claude Code – Claude Sonnet 5.5 (medium) | 65* (attempts per task: 4) | $0.052 | No |
| Claude Code – Claude Sonnet 5.5 (high) | 68* (attempts per task: 4) | $0.067 | No |
| Claude Code – Claude Haiku 5.5 (medium) | 65* (attempts per task: 4) | $0.0074 | Yes |
| Claude Code – Claude Haiku 5.5 (high) | 66* (attempts per task: 3.8) | $0.011 | No |
| Claude Code – Claude Opus 5.5 (xhigh) | 68* (attempts per task: 1.3) | $0.398 | No |
| Claude Code – Claude Sonnet 5.5 (xhigh) | 64* (attempts per task: 1.3) | $0.173 | No |
| Claude Code – Claude Haiku 5.5 (xhigh) | 68* (attempts per task: 1.3) | $0.024 | No |
| Codex – GPT-6.1 Sol (medium) | 70* (attempts per task: 3.3) | $0.060 | Yes |
| Codex – GPT-6.1 Sol (high) | 69* (attempts per task: 3.3) | $0.080 | No |
| Codex – GPT-6.1 Sol (xhigh) | 68* (attempts per task: 3.3) | $0.124 | No |
| Codex – GPT-6 Luna (xhigh) | 69* (attempts per task: 3.3) | $0.0095 | Yes |
Head-to-head comparisons
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – DeepSeek V4 Flash (max)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – DeepSeek V4 Flash (high)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – DeepSeek V4 Pro (high)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – Gemini 3.1 Flash-Lite (low)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – Gemini 3.1 Flash-Lite (minimal)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – Gemini 3.1 Flash-Lite (high)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – Gemini 3.1 Pro (preview) (low)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – Gemini 3.1 Pro (preview) (high)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – Gemini 3.8 Flash (low)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – Gemma 4 26B A4B (thinking on)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – GLM-4.7-Flash (thinking on)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – GLM-5.2 (high)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – GLM-5.2 (none)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – GLM-5.2 (max)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – GLM-5.3-Flash (low)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – GLM-5.3-Flash (max)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – GLM-5.3 (low)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – GLM-5.3 (max, default)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – gpt-oss-120b (low)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – gpt-oss-120b (high)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – Kimi K2.6 (always on)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – Kimi K2.7 Code (always on)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – Mistral Small 3.1 24B (Non-reasoning)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Terminus-2 – Qwen 3.8 27B (xhigh)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Claude Code – Claude Opus 5.5 (medium)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Claude Code – Claude Opus 5.5 (high)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Claude Code – Claude Sonnet 5.5 (medium)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Claude Code – Claude Sonnet 5.5 (high)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Claude Code – Claude Haiku 5.5 (medium)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Claude Code – Claude Haiku 5.5 (high)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Claude Code – Claude Opus 5.5 (xhigh)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Claude Code – Claude Sonnet 5.5 (xhigh)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Claude Code – Claude Haiku 5.5 (xhigh)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Codex – GPT-6.1 Sol (medium)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Codex – GPT-6.1 Sol (high)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Codex – GPT-6.1 Sol (xhigh)
- Terminus-2 – Gemini 3.8 Flash (high) vs. Codex – GPT-6 Luna (xhigh)
Specification and settings
- Developer
- Context window
- 1,048,576 tokens
- Image input
- Yes
- Reasoning setting
- high
- Temperature
- 1.0
- Maximum output tokens
- 65,536
- Input price per 1M tokens
- $0.75
- Cached input price per 1M tokens
- $0.075
- Output price per 1M tokens
- $3.75
Why this model was chosen: model selection.
