Axis Spatial Geospatial Agent Index
The Geospatial Agent Index combines 4 published geospatial benchmarks into one score, weighting each benchmark equally so that a benchmark with many tasks does not outweigh one with few.
Geospatial Agent Index
Composite index of 4 benchmarks, 290 tasks in all:
Each benchmark score is the average pass@1 over its tasks, with three attempts per task and model; the index weights the benchmarks equally. Methodology version 0.4. How we score.
Models7 of 7 models
Score
Axis Spatial Geospatial Agent Index
Equal-weight average of the evaluation scores, 0 to 100 · Higher is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
How to read this chart
What this metric means. The equal-weight average of the 4 evaluation scores. Each evaluation score is the average pass@1 over that benchmark's tasks, from 0 to 100.
Not yet ranked. Entries without a score on every benchmark come after the ranked entries and are not ranked: their index averages only the benchmarks they have.
Partial coverage. A hatched bar, and an asterisk in tables, marks an entry that does not yet have results for every planned attempt. While attempts await review, the table gives the range the score could still reach.
Data table: Axis Spatial Geospatial Agent Index
| Model | Creator | Geospatial Agent Index | Range while attempts are pending | Benchmarks covered | Coverage (attempts) |
|---|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 73* | No pending judgements; coverage incomplete | 4 of 4 | 195 of 870 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 73* | No pending judgements; coverage incomplete | 4 of 4 | 249 of 870 planned attempts |
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 70* | 70 to 70 | 4 of 4 | 169 of 870 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 60* | 60 to 60 | 4 of 4 | 145 of 870 planned attempts | |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 53* | No pending judgements; coverage incomplete | 4 of 4 | 249 of 870 planned attempts |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 43* | 43 to 44 | 4 of 4 | 143 of 870 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 62* | No pending judgements; coverage incomplete | 3 of 4 | 101 of 870 planned attempts |
| Model | GeoAgentBench | GeoBenchX | Earth-Bench | GeoAnalystBench | Geospatial Agent Index |
|---|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Pro (high) | 98 | 45 | 48 | 100 | 73* |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 90 | 54 | 62 | 84 | 73* |
| Terminus-2 – DeepSeek V4 Flash (high) | 96 | 44 | 40 | 100 | 70* |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 86 | 55 | 30 | 68 | 60* |
| Terminus-2 – gpt-oss-120b (high) | 70 | 42 | 46 | 53 | 53* |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 36 | 46 | 48 | 42 | 43* |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 60 | No data | 27 | 100 | 62* |
Cost and time
Axis Spatial Geospatial Agent Index vs. cost per task
Higher and further left is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Most attractive quadrant
- Pareto line
- Terminus-2
- Partial coverage
- Most attractive quadrant
- Pareto line
How to read this chart
What cost is measuring. The average cost of one attempt at a task: tokens reported by the provider multiplied by the serving provider's published list prices, with cached input at its own price. Failed attempts count, so a model that fails slowly and expensively is not flattered.
How to read this chart. Lower is better. Costs are pay-per-token list prices, not any discount or credit a customer might have.
Data table: Axis Spatial Geospatial Agent Index vs. cost per task
| Model | Geospatial Agent Index | Cost per task (USD) | On the Pareto line |
|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | 70* | $0.060 | Yes |
| Terminus-2 – DeepSeek V4 Pro (high) | 73* | $0.261 | Yes |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 60* | $0.012 | Yes |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 43* | $0.025 | No |
| Terminus-2 – gpt-oss-120b (high) | 53* | $0.026 | No |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 73* | $0.126 | Yes |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 62* | $0.066 | No |
Axis Spatial Geospatial Agent Index vs. time per task
Higher and further left is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Most attractive quadrant
- Pareto line
- Terminus-2
- Partial coverage
- Most attractive quadrant
- Pareto line
How to read this chart
What execution time is measuring. Agent wall time per attempt, from the agent's first model call to its last action: model thinking, waiting for the provider and running commands. Sandbox start-up and grading are not included.
How to read this chart. Lower is better. An attempt that reaches the task's time limit scores 0 and counts at the limit.
Data table: Axis Spatial Geospatial Agent Index vs. time per task
| Model | Geospatial Agent Index | Time per task | On the Pareto line |
|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | 70* | 10.0 min | No |
| Terminus-2 – DeepSeek V4 Pro (high) | 73* | 8.6 min | Yes |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 60* | 8.6 min | No |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 43* | 9.5 min | No |
| Terminus-2 – gpt-oss-120b (high) | 53* | 4.6 min | Yes |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 73* | 6.7 min | Yes |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 62* | 19.8 min | No |
Methodology
Each attempt scores 1 or 0. A benchmark score is the average over its tasks of the share of attempts that passed (pass@1). The index is the plain average of the benchmark scores. We run every model ourselves with the same agent, data and limits; the scores are our results on these tasks, not reproductions of the source papers' numbers. Full methodology.
