Terminus-2 – gpt-oss-120b (high): geospatial agent results
Terminus-2 – gpt-oss-120b (high) by OpenAI, run under the Terminus-2 agent: its Geospatial Agent Index, score on each benchmark and each kind of geospatial work, cost, time and token use.
Summary
- Geospatial Agent Index
- 53* (rank 5 of 6)
- Cost per task
- $0.026
- Time per task
- 4.6 min
- Tokens per attempt
- 59k
- Turns per attempt
- 5.3
- Attempts decided
- 249
Score by benchmark
Terminus-2 – gpt-oss-120b (high): score by benchmark
Average pass@1, 0 to 100 · Higher is better
Data table: Terminus-2 – gpt-oss-120b (high): score by benchmark
| Benchmark | Score | Tasks with a decided attempt | Cost per task | Time per task |
|---|---|---|---|---|
| GeoAgentBench | 70 | 50 of 50 | $0.023 | 5.9 min |
| GeoBenchX | 42 | 132 of 173 | $0.022 | 3.7 min |
| Earth-Bench | 46 | 48 of 48 | $0.041 | 5.8 min |
| GeoAnalystBench | 53 | 19 of 19 | $0.018 | 3.6 min |
Score by task format and kind of work
Terminus-2 – gpt-oss-120b (high): score by task format
Average pass@1 over the tasks of each task format, 0 to 100 · Higher is better
How to read this chart
What this shows. The same attempts as the index, grouped by task format. Groups with few tasks are less certain. About these breakdowns.
Data table: Terminus-2 – gpt-oss-120b (high): score by task format
| Task format | Score | Tasks with a decided attempt |
|---|---|---|
| Multi-step analysis | 52 | 128 of 151 |
| Single-answer questions | 46 | 59 of 60 |
| Recognising infeasible requests | 47 | 62 of 79 |
Terminus-2 – gpt-oss-120b (high): score by kind of work
Average pass@1 over the tasks of each kind of work, 0 to 100 · Higher is better
How to read this chart
What this shows. The same attempts as the index, grouped by kind of work. Groups with few tasks are less certain. About these breakdowns.
Data table: Terminus-2 – gpt-oss-120b (high): score by kind of work
| Kind of work | Score | Tasks with a decided attempt |
|---|---|---|
| Raster and remote sensing | 52 | 73 of 75 |
| Vector and overlay | 56 | 18 of 18 |
| Networks and routing | 100 | 5 of 5 |
| Climate and time series | 62 | 24 of 24 |
| Spatial statistics and interpolation | 30 | 20 of 24 |
| Mapping and cartography | 43 | 47 of 65 |
| Recognising infeasible requests | 47 | 62 of 79 |
Compared with other models
Axis Spatial Geospatial Agent Index
Equal-weight average of the evaluation scores, 0 to 100 · Higher is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
How to read this chart
What this metric means. The equal-weight average of the 4 evaluation scores. Each evaluation score is the average pass@1 over that benchmark's tasks, from 0 to 100.
Not yet ranked. Entries without a score on every benchmark come after the ranked entries and are not ranked: their index averages only the benchmarks they have.
Partial coverage. A hatched bar, and an asterisk in tables, marks an entry that does not yet have results for every planned attempt. While attempts await review, the table gives the range the score could still reach.
Data table: Axis Spatial Geospatial Agent Index
| Model | Creator | Geospatial Agent Index | Range while attempts are pending | Benchmarks covered | Coverage (attempts) |
|---|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 73* | No pending judgements; coverage incomplete | 4 of 4 | 195 of 870 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 73* | No pending judgements; coverage incomplete | 4 of 4 | 249 of 870 planned attempts |
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 70* | 70 to 70 | 4 of 4 | 169 of 870 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 60* | 60 to 60 | 4 of 4 | 145 of 870 planned attempts | |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 53* | No pending judgements; coverage incomplete | 4 of 4 | 249 of 870 planned attempts |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 43* | 43 to 44 | 4 of 4 | 143 of 870 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 62* | No pending judgements; coverage incomplete | 3 of 4 | 101 of 870 planned attempts |
Axis Spatial Geospatial Agent Index vs. cost per task
Higher and further left is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Most attractive quadrant
- Pareto line
- Terminus-2
- Partial coverage
- Most attractive quadrant
- Pareto line
Data table: Axis Spatial Geospatial Agent Index vs. cost per task
| Model | Geospatial Agent Index | Cost per task (USD) | On the Pareto line |
|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | 70* | $0.060 | Yes |
| Terminus-2 – DeepSeek V4 Pro (high) | 73* | $0.261 | Yes |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 60* | $0.012 | Yes |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 43* | $0.025 | No |
| Terminus-2 – gpt-oss-120b (high) | 53* | $0.026 | No |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 73* | $0.126 | Yes |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 62* | $0.066 | No |
Head-to-head comparisons
- Terminus-2 – gpt-oss-120b (high) vs. Terminus-2 – DeepSeek V4 Flash (high)
- Terminus-2 – gpt-oss-120b (high) vs. Terminus-2 – DeepSeek V4 Pro (high)
- Terminus-2 – gpt-oss-120b (high) vs. Terminus-2 – Gemma 4 26B A4B (Reasoning)
- Terminus-2 – gpt-oss-120b (high) vs. Terminus-2 – GLM-4.7-Flash (Reasoning)
- Terminus-2 – gpt-oss-120b (high) vs. Terminus-2 – Kimi K2.7 Code (Reasoning)
- Terminus-2 – gpt-oss-120b (high) vs. Terminus-2 – Qwen 3.8 27B (xhigh)
Specification and settings
- Developer
- OpenAI
- Context window
- 128,000 tokens
- Image input
- No
- Reasoning setting
- high
- Temperature
- 0.6
- Maximum output tokens
- 65,536
- Input price per 1M tokens
- $0.35
- Cached input price per 1M tokens
- Not offered
- Output price per 1M tokens
- $0.75
Why this model was chosen: model selection.
