Terminus-2 – GLM-5.3-Flash (low): geospatial agent results
Terminus-2 – GLM-5.3-Flash (low) by Z.ai, run under the Terminus-2 agent: its Geospatial Agent Index, score on each benchmark and each kind of geospatial work, cost, time and token use.
Summary
- Geospatial Agent Index
- 66* (partial: 1.1 of 3 attempts) (rank 25 of 30)
- Cost per task
- $0.028
- Time per task
- 9.8 min
- Tokens per attempt
- 278k
- Turns per attempt
- 12.1
- Attempts decided
- 288
Score by benchmark
Terminus-2 – GLM-5.3-Flash (low): score by benchmark
Average pass@1, 0 to 100 · Higher is better
Data table: Terminus-2 – GLM-5.3-Flash (low): score by benchmark
| Benchmark | Score | Tasks with a decided attempt | Cost per task | Time per task |
|---|---|---|---|---|
| GeoAgentBench | 94 | 49 of 50 | $0.010 | 5.0 min |
| GeoBenchX | 47 | 163 of 173 | $0.028 | 10.0 min |
| Earth-Bench | 27 | 48 of 48 | $0.055 | 16.8 min |
| GeoAnalystBench | 94 | 18 of 19 | $0.010 | 4.6 min |
Score by task format and kind of work
Terminus-2 – GLM-5.3-Flash (low): score by task format
Average pass@1 over the tasks of each task format, 0 to 100 · Higher is better
How to read this chart
What this shows. The same attempts as the index, grouped by task format. Groups with few tasks are less certain. About these breakdowns.
Data table: Terminus-2 – GLM-5.3-Flash (low): score by task format
| Task format | Score | Tasks with a decided attempt |
|---|---|---|
| Multi-step analysis | 75 | 142 of 151 |
| Single-answer questions | 31 | 59 of 60 |
| Recognising infeasible requests | 35 | 77 of 79 |
Terminus-2 – GLM-5.3-Flash (low): score by kind of work
Average pass@1 over the tasks of each kind of work, 0 to 100 · Higher is better
How to read this chart
What this shows. The same attempts as the index, grouped by kind of work. Groups with few tasks are less certain. About these breakdowns.
Data table: Terminus-2 – GLM-5.3-Flash (low): score by kind of work
| Kind of work | Score | Tasks with a decided attempt |
|---|---|---|
| Raster and remote sensing | 61 | 73 of 75 |
| Vector and overlay | 69 | 18 of 18 |
| Networks and routing | 100 | 5 of 5 |
| Climate and time series | 54 | 24 of 24 |
| Spatial statistics and interpolation | 48 | 21 of 24 |
| Mapping and cartography | 67 | 60 of 65 |
| Recognising infeasible requests | 35 | 77 of 79 |
Compared with other models
Geospatial Agent Index
Equal-weight average of the evaluation scores, 0 to 100 · Higher is better
- Partial coverage
How to read this chart
What this metric means. The equal-weight average of the 4 evaluation scores. Each evaluation score is the average pass@1 over that benchmark's tasks, from 0 to 100.
Not yet ranked. Entries without a score on every benchmark come after the ranked entries and are not ranked: their index averages only the benchmarks they have.
Partial coverage. A hatched bar, and an asterisk in tables, marks an entry that does not yet have results for every planned attempt. While attempts await review, the table gives the range the score could still reach.
Partial attempts. Each task is planned 3 times. An entry with fewer attempts per task is in the index with the marker "partial: n of 3 attempts", where n is its attempts per task; its score rests on fewer attempts and can change more.
Data table: Geospatial Agent Index
| Model | Creator | Geospatial Agent Index | Range while attempts are pending | Benchmarks covered | Coverage (attempts) |
|---|---|---|---|---|---|
| Claude Code – Claude Opus 5.5 (high) | Anthropic | 83* (partial: 1.7 of 3 attempts) | 71 to 83 | 4 of 4 | 422 of 870 planned attempts |
| Claude Code – Claude Opus 5.5 (xhigh) | Anthropic | 82 (partial: 1 of 3 attempts) | No pending judgements | 4 of 4 | 290 of 290 planned attempts |
| Claude Code – Claude Opus 5.5 (medium) | Anthropic | 81* (partial: 1.8 of 3 attempts) | 70 to 81 | 4 of 4 | 461 of 870 planned attempts |
| Claude Code – Claude Sonnet 5.5 (high) | Anthropic | 81* (partial: 1.9 of 3 attempts) | 71 to 82 | 4 of 4 | 479 of 870 planned attempts (236 excluded) |
| Codex – GPT-6.1 Sol (high) | OpenAI | 80* (partial: 2.8 of 3 attempts) | 72 to 80 | 4 of 4 | 753 of 870 planned attempts |
| Codex – GPT-6.1 Sol (medium) | OpenAI | 79* (partial: 2.6 of 3 attempts) | No pending judgements; coverage incomplete | 4 of 4 | 749 of 870 planned attempts (6 excluded) |
| Codex – GPT-6.1 Sol (xhigh) | OpenAI | 78* (partial: 2 of 3 attempts) | 74 to 79 | 4 of 4 | 551 of 870 planned attempts |
| Claude Code – Claude Haiku 5.5 (xhigh) | Anthropic | 78 (partial: 1 of 3 attempts) | No pending judgements | 4 of 4 | 290 of 290 planned attempts |
| Codex – GPT-6 Luna (xhigh) | OpenAI | 77* (partial: 1.9 of 3 attempts) | 72 to 78 | 4 of 4 | 512 of 870 planned attempts |
| Claude Code – Claude Haiku 5.5 (medium) | Anthropic | 77* (partial: 1.8 of 3 attempts) | 66 to 78 | 4 of 4 | 462 of 870 planned attempts |
| Claude Code – Claude Sonnet 5.5 (medium) | Anthropic | 77* (partial: 2 of 3 attempts) | 67 to 77 | 4 of 4 | 500 of 870 planned attempts (229 excluded) |
| Claude Code – Claude Haiku 5.5 (high) | Anthropic | 75* (partial: 1.7 of 3 attempts) | 66 to 77 | 4 of 4 | 422 of 870 planned attempts |
| Terminus-2 – GLM-5.2 (max) | Z.ai | 75* (partial: 1.1 of 3 attempts) | 72 to 75 | 4 of 4 | 306 of 870 planned attempts |
| Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high) | Anthropic | 74* (partial: 1.1 of 3 attempts) | 73 to 75 | 4 of 4 | 306 of 870 planned attempts (133 excluded) |
| Terminus-2 – DeepSeek V4 Flash (max) | DeepSeek | 74* (partial: 1.1 of 3 attempts) | 73 to 74 | 4 of 4 | 321 of 870 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 74 | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Claude Code – Claude Sonnet 5.5 (xhigh) | Anthropic | 74 (partial: 1 of 3 attempts) | No pending judgements | 4 of 4 | 290 of 290 planned attempts |
| Terminus-2 – Kimi K2.7 Code (always on) | Moonshot AI | 74 | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Terminus-2 – Kimi K2.6 (always on) | Moonshot AI | 72 | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Terminus-2 – GLM-5.2 (high) | Z.ai | 71* (partial: 1 of 3 attempts) | 65 to 72 | 4 of 4 | 273 of 290 planned attempts |
| Terminus-2 – GLM-5.3 (low) | Z.ai | 71* (partial: 0.9 of 3 attempts) | 61 to 72 | 4 of 4 | 238 of 290 planned attempts |
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 70* | 70 to 71 | 4 of 4 | 860 of 870 planned attempts |
| Terminus-2 – GLM-5.2 (none) | Z.ai | 69* (partial: 1.4 of 3 attempts) | 62 to 72 | 4 of 4 | 356 of 870 planned attempts |
| Terminus-2 – GLM-5.3 (max, default) | Z.ai | 67* (partial: 1.1 of 3 attempts) | 66 to 68 | 4 of 4 | 310 of 870 planned attempts |
| Terminus-2 – GLM-5.3-Flash (low) | Z.ai | 66* (partial: 1.1 of 3 attempts) | 59 to 67 | 4 of 4 | 288 of 870 planned attempts (3 excluded) |
| Terminus-2 – Gemma 4 26B A4B (thinking on) | 62 | No pending judgements | 4 of 4 | 870 of 870 planned attempts | |
| Terminus-2 – GLM-5.3-Flash (max) | Z.ai | 56* (partial: 1.1 of 3 attempts) | 54 to 56 | 4 of 4 | 298 of 870 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 52 | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Terminus-2 – GLM-4.7-Flash (thinking on) | Z.ai | 40 | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Terminus-2 – gpt-oss-120b (low) | OpenAI | 37 | No pending judgements | 4 of 4 | 870 of 870 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 62* (partial: 0.3 of 3 attempts) | No pending judgements; coverage incomplete | 3 of 4 | 101 of 870 planned attempts |
Geospatial Agent Index vs. cost per task
Higher and further left is better
- Partial coverage
- Most attractive quadrant
- Pareto line
Data table: Geospatial Agent Index vs. cost per task
| Model | Geospatial Agent Index | Cost per task (USD) | On the Pareto line |
|---|---|---|---|
| Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high) | 74* (partial: 1.1 of 3 attempts) | $0.278 | No |
| Terminus-2 – DeepSeek V4 Flash (max) | 74* (partial: 1.1 of 3 attempts) | $0.067 | No |
| Terminus-2 – DeepSeek V4 Flash (high) | 70* | $0.067 | No |
| Terminus-2 – DeepSeek V4 Pro (high) | 74 | $0.273 | No |
| Terminus-2 – Gemma 4 26B A4B (thinking on) | 62 | $0.016 | No |
| Terminus-2 – GLM-4.7-Flash (thinking on) | 40 | $0.030 | No |
| Terminus-2 – GLM-5.2 (high) | 71* (partial: 1 of 3 attempts) | $0.239 | No |
| Terminus-2 – GLM-5.2 (none) | 69* (partial: 1.4 of 3 attempts) | $0.103 | No |
| Terminus-2 – GLM-5.2 (max) | 75* (partial: 1.1 of 3 attempts) | $0.362 | No |
| Terminus-2 – GLM-5.3-Flash (low) | 66* (partial: 1.1 of 3 attempts) | $0.028 | No |
| Terminus-2 – GLM-5.3-Flash (max) | 56* (partial: 1.1 of 3 attempts) | $0.042 | No |
| Terminus-2 – GLM-5.3 (low) | 71* (partial: 0.9 of 3 attempts) | $0.250 | No |
| Terminus-2 – GLM-5.3 (max, default) | 67* (partial: 1.1 of 3 attempts) | $0.363 | No |
| Terminus-2 – gpt-oss-120b (low) | 37 | $0.0088 | No |
| Terminus-2 – gpt-oss-120b (high) | 52 | $0.026 | No |
| Terminus-2 – Kimi K2.6 (always on) | 72 | $0.072 | No |
| Terminus-2 – Kimi K2.7 Code (always on) | 74 | $0.119 | No |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 62* (partial: 0.3 of 3 attempts) | $0.066 | No |
| Claude Code – Claude Opus 5.5 (medium) | 81* (partial: 1.8 of 3 attempts) | $0.138 | Yes |
| Claude Code – Claude Opus 5.5 (high) | 83* (partial: 1.7 of 3 attempts) | $0.175 | Yes |
| Claude Code – Claude Sonnet 5.5 (medium) | 77* (partial: 2 of 3 attempts) | $0.052 | No |
| Claude Code – Claude Sonnet 5.5 (high) | 81* (partial: 1.9 of 3 attempts) | $0.062 | Yes |
| Claude Code – Claude Haiku 5.5 (medium) | 77* (partial: 1.8 of 3 attempts) | $0.0075 | Yes |
| Claude Code – Claude Haiku 5.5 (high) | 75* (partial: 1.7 of 3 attempts) | $0.010 | No |
| Claude Code – Claude Opus 5.5 (xhigh) | 82 (partial: 1 of 3 attempts) | $0.375 | No |
| Claude Code – Claude Sonnet 5.5 (xhigh) | 74 (partial: 1 of 3 attempts) | $0.163 | No |
| Claude Code – Claude Haiku 5.5 (xhigh) | 78 (partial: 1 of 3 attempts) | $0.023 | Yes |
| Codex – GPT-6.1 Sol (medium) | 79* (partial: 2.6 of 3 attempts) | $0.060 | Yes |
| Codex – GPT-6.1 Sol (high) | 80* (partial: 2.8 of 3 attempts) | $0.080 | No |
| Codex – GPT-6.1 Sol (xhigh) | 78* (partial: 2 of 3 attempts) | $0.124 | No |
| Codex – GPT-6 Luna (xhigh) | 77* (partial: 1.9 of 3 attempts) | $0.0092 | Yes |
Head-to-head comparisons
- Terminus-2 – GLM-5.3-Flash (low) vs. Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high)
- Terminus-2 – GLM-5.3-Flash (low) vs. Terminus-2 – DeepSeek V4 Flash (max)
- Terminus-2 – GLM-5.3-Flash (low) vs. Terminus-2 – DeepSeek V4 Flash (high)
- Terminus-2 – GLM-5.3-Flash (low) vs. Terminus-2 – DeepSeek V4 Pro (high)
- Terminus-2 – GLM-5.3-Flash (low) vs. Terminus-2 – Gemma 4 26B A4B (thinking on)
- Terminus-2 – GLM-5.3-Flash (low) vs. Terminus-2 – GLM-4.7-Flash (thinking on)
- Terminus-2 – GLM-5.3-Flash (low) vs. Terminus-2 – GLM-5.2 (high)
- Terminus-2 – GLM-5.3-Flash (low) vs. Terminus-2 – GLM-5.2 (none)
- Terminus-2 – GLM-5.3-Flash (low) vs. Terminus-2 – GLM-5.2 (max)
- Terminus-2 – GLM-5.3-Flash (low) vs. Terminus-2 – GLM-5.3-Flash (max)
- Terminus-2 – GLM-5.3-Flash (low) vs. Terminus-2 – GLM-5.3 (low)
- Terminus-2 – GLM-5.3-Flash (low) vs. Terminus-2 – GLM-5.3 (max, default)
- Terminus-2 – GLM-5.3-Flash (low) vs. Terminus-2 – gpt-oss-120b (low)
- Terminus-2 – GLM-5.3-Flash (low) vs. Terminus-2 – gpt-oss-120b (high)
- Terminus-2 – GLM-5.3-Flash (low) vs. Terminus-2 – Kimi K2.6 (always on)
- Terminus-2 – GLM-5.3-Flash (low) vs. Terminus-2 – Kimi K2.7 Code (always on)
- Terminus-2 – GLM-5.3-Flash (low) vs. Terminus-2 – Qwen 3.8 27B (xhigh)
- Terminus-2 – GLM-5.3-Flash (low) vs. Claude Code – Claude Opus 5.5 (medium)
- Terminus-2 – GLM-5.3-Flash (low) vs. Claude Code – Claude Opus 5.5 (high)
- Terminus-2 – GLM-5.3-Flash (low) vs. Claude Code – Claude Sonnet 5.5 (medium)
- Terminus-2 – GLM-5.3-Flash (low) vs. Claude Code – Claude Sonnet 5.5 (high)
- Terminus-2 – GLM-5.3-Flash (low) vs. Claude Code – Claude Haiku 5.5 (medium)
- Terminus-2 – GLM-5.3-Flash (low) vs. Claude Code – Claude Haiku 5.5 (high)
- Terminus-2 – GLM-5.3-Flash (low) vs. Claude Code – Claude Opus 5.5 (xhigh)
- Terminus-2 – GLM-5.3-Flash (low) vs. Claude Code – Claude Sonnet 5.5 (xhigh)
- Terminus-2 – GLM-5.3-Flash (low) vs. Claude Code – Claude Haiku 5.5 (xhigh)
- Terminus-2 – GLM-5.3-Flash (low) vs. Codex – GPT-6.1 Sol (medium)
- Terminus-2 – GLM-5.3-Flash (low) vs. Codex – GPT-6.1 Sol (high)
- Terminus-2 – GLM-5.3-Flash (low) vs. Codex – GPT-6.1 Sol (xhigh)
- Terminus-2 – GLM-5.3-Flash (low) vs. Codex – GPT-6 Luna (xhigh)
Specification and settings
- Developer
- Z.ai
- Context window
- 1,048,576 tokens
- Image input
- Yes
- Reasoning setting
- low
- Temperature
- 0.6
- Maximum output tokens
- 65,536
- Input price per 1M tokens
- $0.15
- Cached input price per 1M tokens
- $0.030
- Output price per 1M tokens
- $0.50
Why this model was chosen: model selection.
