Alibaba (Qwen) models: geospatial agent results
Geospatial Agent Benchmarks tests 1 model configuration from Alibaba (Qwen).
Highlights
- Highest Geospatial Agent Index (ranked entries): No data yet.
- Lowest cost per task: Terminus-2 – Qwen 3.8 27B (xhigh) ($0.066).
- Fastest per task: Terminus-2 – Qwen 3.8 27B (xhigh) (19.8 min).
Results
Axis Spatial Geospatial Agent Index
Equal-weight average of the evaluation scores, 0 to 100 · Higher is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
Data table: Axis Spatial Geospatial Agent Index
| Model | Creator | Geospatial Agent Index | Range while attempts are pending | Benchmarks covered | Coverage (attempts) |
|---|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 73* | No pending judgements; coverage incomplete | 4 of 4 | 195 of 870 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 73* | No pending judgements; coverage incomplete | 4 of 4 | 249 of 870 planned attempts |
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 70* | 70 to 70 | 4 of 4 | 169 of 870 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 60* | 60 to 60 | 4 of 4 | 145 of 870 planned attempts | |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 53* | No pending judgements; coverage incomplete | 4 of 4 | 249 of 870 planned attempts |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 43* | 43 to 44 | 4 of 4 | 143 of 870 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 62* | No pending judgements; coverage incomplete | 3 of 4 | 101 of 870 planned attempts |
| Model | Geospatial Agent Index | GeoAgentBench | GeoBenchX | Earth-Bench | GeoAnalystBench | Cost per task | Time per task | Context window |
|---|---|---|---|---|---|---|---|---|
| Terminus-2 – Qwen 3.8 27B (xhigh) | 62* | 60 | No data | 27 | 100 | $0.066 | 19.8 min | 262,144 |
