Terminus-2 – Kimi K2.7 Code (always on): geospatial agent results
Terminus-2 – Kimi K2.7 Code (always on) by Moonshot AI, run under the Terminus-2 agent: its Geospatial Agent Index, score on each benchmark and each kind of geospatial work, cost, time and token use.
Summary
- Geospatial Agent Index
- 73* (rank 2 of 6)
- Cost per task
- $0.129
- Time per task
- 6.8 min
- Tokens per attempt
- 308k
- Turns per attempt
- 14.6
- Attempts decided
- 268
Score by benchmark
Terminus-2 – Kimi K2.7 Code (always on): score by benchmark
Average pass@1, 0 to 100 · Higher is better
Data table: Terminus-2 – Kimi K2.7 Code (always on): score by benchmark
| Benchmark | Score | Tasks with a decided attempt | Cost per task | Time per task |
|---|---|---|---|---|
| GeoAgentBench | 90 | 50 of 50 | $0.038 | 3.0 min |
| GeoBenchX | 55 | 147 of 173 | $0.114 | 6.4 min |
| Earth-Bench | 62 | 48 of 48 | $0.288 | 12.3 min |
| GeoAnalystBench | 84 | 19 of 19 | $0.040 | 4.1 min |
Score by task format and kind of work
Terminus-2 – Kimi K2.7 Code (always on): score by task format
Average pass@1 over the tasks of each task format, 0 to 100 · Higher is better
How to read this chart
What this shows. The same attempts as the index, grouped by task format. Groups with few tasks are less certain. About these breakdowns.
Data table: Terminus-2 – Kimi K2.7 Code (always on): score by task format
| Task format | Score | Tasks with a decided attempt |
|---|---|---|
| Multi-step analysis | 80 | 133 of 151 |
| Single-answer questions | 59 | 60 of 60 |
| Recognising infeasible requests | 41 | 71 of 79 |
Terminus-2 – Kimi K2.7 Code (always on): score by kind of work
Average pass@1 over the tasks of each kind of work, 0 to 100 · Higher is better
How to read this chart
What this shows. The same attempts as the index, grouped by kind of work. Groups with few tasks are less certain. About these breakdowns.
Data table: Terminus-2 – Kimi K2.7 Code (always on): score by kind of work
| Kind of work | Score | Tasks with a decided attempt |
|---|---|---|
| Raster and remote sensing | 74 | 74 of 75 |
| Vector and overlay | 78 | 18 of 18 |
| Networks and routing | 80 | 5 of 5 |
| Climate and time series | 79 | 24 of 24 |
| Spatial statistics and interpolation | 67 | 21 of 24 |
| Mapping and cartography | 73 | 51 of 65 |
| Recognising infeasible requests | 41 | 71 of 79 |
Compared with other models
Axis Spatial Geospatial Agent Index
Equal-weight average of the evaluation scores, 0 to 100 · Higher is better
- Partial coverage
How to read this chart
What this metric means. The equal-weight average of the 4 evaluation scores. Each evaluation score is the average pass@1 over that benchmark's tasks, from 0 to 100.
Not yet ranked. Entries without a score on every benchmark come after the ranked entries and are not ranked: their index averages only the benchmarks they have.
Partial coverage. A hatched bar, and an asterisk in tables, marks an entry that does not yet have results for every planned attempt. While attempts await review, the table gives the range the score could still reach.
Data table: Axis Spatial Geospatial Agent Index
| Model | Creator | Geospatial Agent Index | Range while attempts are pending | Benchmarks covered | Coverage (attempts) |
|---|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 73* | No pending judgements; coverage incomplete | 4 of 4 | 209 of 870 planned attempts |
| Terminus-2 – Kimi K2.7 Code (always on) | Moonshot AI | 73* | 73 to 73 | 4 of 4 | 268 of 870 planned attempts |
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 71* | No pending judgements; coverage incomplete | 4 of 4 | 181 of 870 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (thinking on) | 61* | 61 to 61 | 4 of 4 | 209 of 870 planned attempts | |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 52* | 52 to 52 | 4 of 4 | 320 of 870 planned attempts |
| Terminus-2 – GLM-4.7-Flash (thinking on) | Z.ai | 41* | No pending judgements; coverage incomplete | 4 of 4 | 176 of 870 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 62* | No pending judgements; coverage incomplete | 3 of 4 | 101 of 870 planned attempts |
Axis Spatial Geospatial Agent Index vs. cost per task
Higher and further left is better
- Partial coverage
- Most attractive quadrant
- Pareto line
Data table: Axis Spatial Geospatial Agent Index vs. cost per task
| Model | Geospatial Agent Index | Cost per task (USD) | On the Pareto line |
|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | 71* | $0.059 | Yes |
| Terminus-2 – DeepSeek V4 Pro (high) | 73* | $0.265 | Yes |
| Terminus-2 – Gemma 4 26B A4B (thinking on) | 61* | $0.014 | Yes |
| Terminus-2 – GLM-4.7-Flash (thinking on) | 41* | $0.027 | No |
| Terminus-2 – gpt-oss-120b (high) | 52* | $0.025 | No |
| Terminus-2 – Kimi K2.7 Code (always on) | 73* | $0.129 | Yes |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 62* | $0.066 | No |
Head-to-head comparisons
- Terminus-2 – Kimi K2.7 Code (always on) vs. Terminus-2 – DeepSeek V4 Flash (high)
- Terminus-2 – Kimi K2.7 Code (always on) vs. Terminus-2 – DeepSeek V4 Pro (high)
- Terminus-2 – Kimi K2.7 Code (always on) vs. Terminus-2 – Gemma 4 26B A4B (thinking on)
- Terminus-2 – Kimi K2.7 Code (always on) vs. Terminus-2 – GLM-4.7-Flash (thinking on)
- Terminus-2 – Kimi K2.7 Code (always on) vs. Terminus-2 – gpt-oss-120b (high)
- Terminus-2 – Kimi K2.7 Code (always on) vs. Terminus-2 – Qwen 3.8 27B (xhigh)
Specification and settings
- Developer
- Moonshot AI
- Context window
- 262,144 tokens
- Image input
- Yes
- Reasoning setting
- always on
- Temperature
- 0.6
- Maximum output tokens
- 65,536
- Input price per 1M tokens
- $0.95
- Cached input price per 1M tokens
- $0.190
- Output price per 1M tokens
- $4.00
Why this model was chosen: model selection.
