Terminus-2 – Gemini 3.1 Pro (preview) (low): geospatial agent results
Terminus-2 – Gemini 3.1 Pro (preview) (low) by Google, run under the Terminus-2 agent: its Geospatial Agent Index, score on each benchmark and each kind of geospatial work, cost, time and token use.
Summary
- Geospatial Agent Index
- 65 (attempts per task: 3) (rank 27 of 35)
- Cost per task
- $0.061
- Time per task
- 1.0 min
- Tokens per attempt
- 19k
- Turns per attempt
- 3.8
- Attempts decided
- 870
Score by benchmark
Terminus-2 – Gemini 3.1 Pro (preview) (low): score by benchmark
Average pass@1, 0 to 100 · Higher is better
Data table: Terminus-2 – Gemini 3.1 Pro (preview) (low): score by benchmark
| Benchmark | Score | Tasks with a decided attempt | Cost per task | Time per task |
|---|---|---|---|---|
| GeoAgentBench | 83 | 50 of 50 | $0.062 | 1.1 min |
| GeoBenchX | 61 | 173 of 173 | $0.058 | 0.8 min |
| Earth-Bench | 42 | 48 of 48 | $0.075 | 1.2 min |
| GeoAnalystBench | 75 | 19 of 19 | $0.057 | 1.1 min |
Score by task format and kind of work
Terminus-2 – Gemini 3.1 Pro (preview) (low): score by task format
Average pass@1 over the tasks of each task format, 0 to 100 · Higher is better
How to read this chart
What this shows. The same attempts as the index, grouped by task format. Groups with few tasks are less certain. About these breakdowns.
Data table: Terminus-2 – Gemini 3.1 Pro (preview) (low): score by task format
| Task format | Score | Tasks with a decided attempt |
|---|---|---|
| Multi-step analysis | 65 | 151 of 151 |
| Single-answer questions | 45 | 60 of 60 |
| Recognising infeasible requests | 71 | 79 of 79 |
Terminus-2 – Gemini 3.1 Pro (preview) (low): score by kind of work
Average pass@1 over the tasks of each kind of work, 0 to 100 · Higher is better
How to read this chart
What this shows. The same attempts as the index, grouped by kind of work. Groups with few tasks are less certain. About these breakdowns.
Data table: Terminus-2 – Gemini 3.1 Pro (preview) (low): score by kind of work
| Kind of work | Score | Tasks with a decided attempt |
|---|---|---|
| Raster and remote sensing | 58 | 75 of 75 |
| Vector and overlay | 69 | 18 of 18 |
| Networks and routing | 93 | 5 of 5 |
| Climate and time series | 67 | 24 of 24 |
| Spatial statistics and interpolation | 51 | 24 of 24 |
| Mapping and cartography | 55 | 65 of 65 |
| Recognising infeasible requests | 71 | 79 of 79 |
Terminus-2 – Gemini 3.1 Pro (preview) (low): score by data type
Average pass@1 over the tasks of each data type, 0 to 100 · Higher is better
How to read this chart
What this shows. The same attempts as the index, grouped by data type. Groups with few tasks are less certain. About these breakdowns.
Data table: Terminus-2 – Gemini 3.1 Pro (preview) (low): score by data type
| Data type | Score | Tasks with a decided attempt |
|---|---|---|
| Optical multispectral | No data | 0 of 43 |
| Elevation or terrain | No data | 0 of 16 |
| Climate or gridded time series | No data | 0 of 17 |
| Other gridded data | No data | 0 of 26 |
| Vector only | No data | 0 of 108 |
| Tabular only | No data | 0 of 2 |
Terminus-2 – Gemini 3.1 Pro (preview) (low): score by task length
Average pass@1 over the tasks of each task length, 0 to 100 · Higher is better
How to read this chart
What this shows. The same attempts as the index, grouped by task length. Groups with few tasks are less certain. About these breakdowns.
Data table: Terminus-2 – Gemini 3.1 Pro (preview) (low): score by task length
| Task length | Score | Tasks with a decided attempt |
|---|
Compared with other models
Geospatial Agent Index
Equal-weight average of the evaluation scores, 0 to 100 · Higher is better
- Partial coverage
How to read this chart
What this metric means. The equal-weight average of the 4 evaluation scores. Each evaluation score is the average pass@1 over that benchmark's tasks, from 0 to 100.
Not yet ranked. Entries without a score on every benchmark come after the ranked entries and are not ranked: their index averages only the benchmarks they have.
Partial coverage. A hatched bar, and an asterisk in tables, marks an entry that does not yet have results for every planned attempt. While attempts await review, the table gives the range the score could still reach.
Attempts per task. Each task's score is the average of its recorded attempts (up to 3). Some entries and tracks currently have 1 attempt per task; further attempts will be added and averaged in. Each entry shows its attempts per task.
Data table: Geospatial Agent Index
| Model | Creator | Geospatial Agent Index | Range while attempts are pending | Benchmarks covered | Coverage (attempts) |
|---|---|---|---|---|---|
| Claude Code – Claude Opus 5.5 (xhigh) | Anthropic | 82 (attempts per task: 1) | No pending judgements | 4 of 4 | 290 of 290 planned attempts |
| Claude Code – Claude Opus 5.5 (high) | Anthropic | 82* (attempts per task: 2.5) | 63 to 84 | 4 of 4 | 531 of 870 planned attempts |
| Claude Code – Claude Sonnet 5.5 (high) | Anthropic | 81* (attempts per task: 2) | 77 to 81 | 4 of 4 | 556 of 870 planned attempts |
| Claude Code – Claude Opus 5.5 (medium) | Anthropic | 80* (attempts per task: 2.7) | 58 to 83 | 4 of 4 | 573 of 870 planned attempts |
| Codex – GPT-6.1 Sol (high) | OpenAI | 80 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Codex – GPT-6.1 Sol (medium) | OpenAI | 79* (attempts per task: 2.6) | No pending judgements; coverage incomplete | 4 of 4 | 749 of 870 planned attempts (6 excluded) |
| Codex – GPT-6.1 Sol (xhigh) | OpenAI | 78* (attempts per task: 2.6) | 69 to 78 | 4 of 4 | 668 of 870 planned attempts |
| Claude Code – Claude Haiku 5.5 (xhigh) | Anthropic | 78 (attempts per task: 1) | No pending judgements | 4 of 4 | 290 of 290 planned attempts |
| Codex – GPT-6 Luna (xhigh) | OpenAI | 77* (attempts per task: 2.4) | 69 to 78 | 4 of 4 | 627 of 870 planned attempts |
| Claude Code – Claude Haiku 5.5 (medium) | Anthropic | 77* (attempts per task: 2) | 74 to 78 | 4 of 4 | 538 of 870 planned attempts |
| Claude Code – Claude Haiku 5.5 (high) | Anthropic | 77* (attempts per task: 2.3) | 61 to 79 | 4 of 4 | 526 of 870 planned attempts |
| Claude Code – Claude Sonnet 5.5 (medium) | Anthropic | 76* (attempts per task: 2.2) | 71 to 76 | 4 of 4 | 580 of 870 planned attempts |
| Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high) | Anthropic | 75* (attempts per task: 1.2) | 74 to 75 | 4 of 4 | 335 of 870 planned attempts (138 excluded) |
| Terminus-2 – Gemini 3.1 Pro (preview) (high) | 75 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts | |
| Terminus-2 – DeepSeek V4 Flash (max) | DeepSeek | 74* (attempts per task: 1.2) | 73 to 75 | 4 of 4 | 344 of 870 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 74 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Claude Code – Claude Sonnet 5.5 (xhigh) | Anthropic | 74 (attempts per task: 1) | No pending judgements | 4 of 4 | 290 of 290 planned attempts |
| Terminus-2 – GLM-5.2 (max) | Z.ai | 74* (attempts per task: 1.2) | 74 to 74 | 4 of 4 | 338 of 870 planned attempts (1 excluded) |
| Terminus-2 – Kimi K2.7 Code (always on) | Moonshot AI | 74 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Terminus-2 – GLM-5.2 (high) | Z.ai | 72* (attempts per task: 1.1) | 71 to 72 | 4 of 4 | 310 of 870 planned attempts (1 excluded) |
| Terminus-2 – Kimi K2.6 (always on) | Moonshot AI | 72 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Terminus-2 – GLM-5.3 (low) | Z.ai | 72* (attempts per task: 1) | 70 to 72 | 4 of 4 | 285 of 290 planned attempts |
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 71 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Terminus-2 – GLM-5.2 (none) | Z.ai | 70* (attempts per task: 1.6) | 68 to 70 | 4 of 4 | 446 of 870 planned attempts |
| Terminus-2 – GLM-5.3 (max, default) | Z.ai | 68* (attempts per task: 1.2) | 67 to 68 | 4 of 4 | 335 of 870 planned attempts |
| Terminus-2 – GLM-5.3-Flash (low) | Z.ai | 66* (attempts per task: 1.4) | 61 to 67 | 4 of 4 | 380 of 870 planned attempts (4 excluded) |
| Terminus-2 – Gemini 3.1 Pro (preview) (low) | 65 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts (22 excluded) | |
| Terminus-2 – Gemma 4 26B A4B (thinking on) | 62 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts | |
| Terminus-2 – Gemini 3.1 Flash-Lite (high) | 59* (attempts per task: 3) | 59 to 59 | 4 of 4 | 864 of 870 planned attempts (6 excluded) | |
| Terminus-2 – GLM-5.3-Flash (max) | Z.ai | 56* (attempts per task: 1.2) | 55 to 56 | 4 of 4 | 342 of 870 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 52 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Terminus-2 – Gemini 3.1 Flash-Lite (low) | 49 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (thinking on) | Z.ai | 40 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Terminus-2 – gpt-oss-120b (low) | OpenAI | 37 (attempts per task: 3) | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Terminus-2 – Gemini 3.1 Flash-Lite (minimal) | 34* (attempts per task: 3) | No pending judgements; coverage incomplete | 4 of 4 | 869 of 870 planned attempts | |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 62* (attempts per task: 1 on 101 of 290 tasks so far) | No pending judgements; coverage incomplete | 3 of 4 | 101 of 870 planned attempts |
| Terminus-2 – Mistral Small 3.1 24B (Non-reasoning) | Mistral AI | 0* (attempts per task: 1 on 9 of 290 tasks so far) | No pending judgements; coverage incomplete | 1 of 4 | 9 of 351 planned attempts |
Geospatial Agent Index vs. cost per task
Higher and further left is better
- Partial coverage
- Most attractive quadrant
- Pareto line
Data table: Geospatial Agent Index vs. cost per task
| Model | Geospatial Agent Index | Cost per task (USD) | On the Pareto line |
|---|---|---|---|
| Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high) | 75* (attempts per task: 1.2) | $0.272 | No |
| Terminus-2 – DeepSeek V4 Flash (max) | 74* (attempts per task: 1.2) | $0.067 | No |
| Terminus-2 – DeepSeek V4 Flash (high) | 71 (attempts per task: 3) | $0.067 | No |
| Terminus-2 – DeepSeek V4 Pro (high) | 74 (attempts per task: 3) | $0.273 | No |
| Terminus-2 – Gemini 3.1 Flash-Lite (low) | 49 (attempts per task: 3) | $0.012 | No |
| Terminus-2 – Gemini 3.1 Flash-Lite (minimal) | 34* (attempts per task: 3) | $0.0056 | Yes |
| Terminus-2 – Gemini 3.1 Flash-Lite (high) | 59* (attempts per task: 3) | $0.021 | No |
| Terminus-2 – Gemini 3.1 Pro (preview) (low) | 65 (attempts per task: 3) | $0.061 | No |
| Terminus-2 – Gemini 3.1 Pro (preview) (high) | 75 (attempts per task: 3) | $0.252 | No |
| Terminus-2 – Gemma 4 26B A4B (thinking on) | 62 (attempts per task: 3) | $0.016 | No |
| Terminus-2 – GLM-4.7-Flash (thinking on) | 40 (attempts per task: 3) | $0.030 | No |
| Terminus-2 – GLM-5.2 (high) | 72* (attempts per task: 1.1) | $0.231 | No |
| Terminus-2 – GLM-5.2 (none) | 70* (attempts per task: 1.6) | $0.102 | No |
| Terminus-2 – GLM-5.2 (max) | 74* (attempts per task: 1.2) | $0.359 | No |
| Terminus-2 – GLM-5.3-Flash (low) | 66* (attempts per task: 1.4) | $0.028 | No |
| Terminus-2 – GLM-5.3-Flash (max) | 56* (attempts per task: 1.2) | $0.042 | No |
| Terminus-2 – GLM-5.3 (low) | 72* (attempts per task: 1) | $0.251 | No |
| Terminus-2 – GLM-5.3 (max, default) | 68* (attempts per task: 1.2) | $0.360 | No |
| Terminus-2 – gpt-oss-120b (low) | 37 (attempts per task: 3) | $0.0088 | No |
| Terminus-2 – gpt-oss-120b (high) | 52 (attempts per task: 3) | $0.026 | No |
| Terminus-2 – Kimi K2.6 (always on) | 72 (attempts per task: 3) | $0.072 | No |
| Terminus-2 – Kimi K2.7 Code (always on) | 74 (attempts per task: 3) | $0.119 | No |
| Terminus-2 – Mistral Small 3.1 24B (Non-reasoning) | 0* (attempts per task: 1 on 9 of 290 tasks so far) | $0.247 | No |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 62* (attempts per task: 1 on 101 of 290 tasks so far) | $0.066 | No |
| Claude Code – Claude Opus 5.5 (medium) | 80* (attempts per task: 2.7) | $0.137 | No |
| Claude Code – Claude Opus 5.5 (high) | 82* (attempts per task: 2.5) | $0.167 | Yes |
| Claude Code – Claude Sonnet 5.5 (medium) | 76* (attempts per task: 2.2) | $0.053 | No |
| Claude Code – Claude Sonnet 5.5 (high) | 81* (attempts per task: 2) | $0.062 | Yes |
| Claude Code – Claude Haiku 5.5 (medium) | 77* (attempts per task: 2) | $0.0074 | Yes |
| Claude Code – Claude Haiku 5.5 (high) | 77* (attempts per task: 2.3) | $0.010 | No |
| Claude Code – Claude Opus 5.5 (xhigh) | 82 (attempts per task: 1) | $0.375 | Yes |
| Claude Code – Claude Sonnet 5.5 (xhigh) | 74 (attempts per task: 1) | $0.163 | No |
| Claude Code – Claude Haiku 5.5 (xhigh) | 78 (attempts per task: 1) | $0.023 | Yes |
| Codex – GPT-6.1 Sol (medium) | 79* (attempts per task: 2.6) | $0.060 | Yes |
| Codex – GPT-6.1 Sol (high) | 80 (attempts per task: 3) | $0.080 | No |
| Codex – GPT-6.1 Sol (xhigh) | 78* (attempts per task: 2.6) | $0.121 | No |
| Codex – GPT-6 Luna (xhigh) | 77* (attempts per task: 2.4) | $0.0092 | Yes |
Head-to-head comparisons
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – DeepSeek V4 Flash (max)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – DeepSeek V4 Flash (high)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – DeepSeek V4 Pro (high)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – Gemini 3.1 Flash-Lite (low)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – Gemini 3.1 Flash-Lite (minimal)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – Gemini 3.1 Flash-Lite (high)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – Gemini 3.1 Pro (preview) (high)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – Gemma 4 26B A4B (thinking on)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – GLM-4.7-Flash (thinking on)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – GLM-5.2 (high)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – GLM-5.2 (none)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – GLM-5.2 (max)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – GLM-5.3-Flash (low)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – GLM-5.3-Flash (max)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – GLM-5.3 (low)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – GLM-5.3 (max, default)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – gpt-oss-120b (low)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – gpt-oss-120b (high)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – Kimi K2.6 (always on)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – Kimi K2.7 Code (always on)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – Mistral Small 3.1 24B (Non-reasoning)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Terminus-2 – Qwen 3.8 27B (xhigh)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Claude Code – Claude Opus 5.5 (medium)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Claude Code – Claude Opus 5.5 (high)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Claude Code – Claude Sonnet 5.5 (medium)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Claude Code – Claude Sonnet 5.5 (high)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Claude Code – Claude Haiku 5.5 (medium)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Claude Code – Claude Haiku 5.5 (high)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Claude Code – Claude Opus 5.5 (xhigh)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Claude Code – Claude Sonnet 5.5 (xhigh)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Claude Code – Claude Haiku 5.5 (xhigh)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Codex – GPT-6.1 Sol (medium)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Codex – GPT-6.1 Sol (high)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Codex – GPT-6.1 Sol (xhigh)
- Terminus-2 – Gemini 3.1 Pro (preview) (low) vs. Codex – GPT-6 Luna (xhigh)
Specification and settings
- Developer
- Context window
- 1,048,576 tokens
- Image input
- Yes
- Reasoning setting
- low
- Temperature
- 1.0
- Maximum output tokens
- 65,536
- Input price per 1M tokens
- $2.00
- Cached input price per 1M tokens
- $0.200
- Output price per 1M tokens
- $12.00
Why this model was chosen: model selection.
