GeoAgentBench benchmark leaderboard
Multi-step spatial execution: 50 tasks from GeoAgentBench, run by Axis Spatial with the same agent and limits for every model. One of the 4 benchmarks in the Geospatial Agent Index.
Models7 of 7 models
Score
GeoAgentBench score
Average pass@1 over its tasks, 0 to 100 · Higher is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
How to read this chart
What this metric means. For each task, the share of a model's attempts that passed every check; then the average over the tasks. Tasks with no decided attempt yet are left out.
Data table: GeoAgentBench score
| Model | Creator | GeoAgentBench score | Range while attempts are pending | Coverage (attempts) |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 96 | No pending judgements; coverage incomplete | 50 of 150 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 98 | No pending judgements; coverage incomplete | 50 of 150 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 86 | No pending judgements; coverage incomplete | 49 of 150 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 36 | No pending judgements; coverage incomplete | 50 of 150 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 70 | No pending judgements; coverage incomplete | 50 of 150 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 90 | No pending judgements; coverage incomplete | 50 of 150 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 60 | No pending judgements; coverage incomplete | 48 of 150 planned attempts |
| Model | Score | Tasks with a decided attempt | Attempts decided | Passed | Time limit reached |
|---|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | 96 | 50 of 50 | 50 | 48 | 0 |
| Terminus-2 – DeepSeek V4 Pro (high) | 98 | 50 of 50 | 50 | 49 | 0 |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 86 | 49 of 50 | 49 | 42 | 1 |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 36 | 50 of 50 | 50 | 18 | 10 |
| Terminus-2 – gpt-oss-120b (high) | 70 | 50 of 50 | 50 | 35 | 2 |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 90 | 50 of 50 | 50 | 45 | 0 |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 60 | 48 of 50 | 48 | 29 | 19 |
Token usage
GeoAgentBench: token usage per attempt
Average tokens per attempt · Lower is better
- Input (not cached)
- Cached input
- Output, including reasoning
- Partial coverage
How to read this chart
What this metric means. Average tokens per attempt, as reported by the provider: input the model read (split into new and cached input) and output it wrote, which includes its reasoning.
Compare with care. Models use different tokenisers, so the same text is a different number of tokens for different models. Use token counts to see how much work each model did, and cost to compare spending.
Data table: GeoAgentBench: token usage per attempt
| Model | Input (not cached) | Cached input | Output, including reasoning | Total | Coverage (attempts) |
|---|---|---|---|---|---|
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 293k | No data | 14k | 307k | 50 of 150 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 97k | No data | 8k | 105k | 48 of 150 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 52k | 24k | 13k | 89k | 49 of 150 planned attempts |
| Terminus-2 – DeepSeek V4 Flash (high) | 11k | 50k | 16k | 77k | 50 of 150 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 7k | 50k | 6k | 63k | 50 of 150 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | 40k | 6k | 10k | 56k | 50 of 150 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | 35k | No data | 14k | 49k | 50 of 150 planned attempts |
Cost
GeoAgentBench: cost per task
Average cost per attempt (USD) · Lower is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
How to read this chart
What cost is measuring. The average cost of one attempt at a task: tokens reported by the provider multiplied by the serving provider's published list prices, with cached input at its own price. Failed attempts count, so a model that fails slowly and expensively is not flattered.
How to read this chart. Lower is better. Costs are pay-per-token list prices, not any discount or credit a customer might have.
Data table: GeoAgentBench: cost per task
| Model | Creator | GeoAgentBench: cost per task | Coverage (attempts) |
|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | $0.026 | 50 of 150 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | $0.091 | 50 of 150 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | $0.012 | 49 of 150 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | $0.023 | 50 of 150 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | $0.023 | 50 of 150 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | $0.038 | 50 of 150 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | $0.070 | 48 of 150 planned attempts |
Time
GeoAgentBench: time per task
Average agent wall time per attempt · Lower is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
How to read this chart
What execution time is measuring. Agent wall time per attempt, from the agent's first model call to its last action: model thinking, waiting for the provider and running commands. Sandbox start-up and grading are not included.
How to read this chart. Lower is better. An attempt that reaches the task's time limit scores 0 and counts at the limit.
Data table: GeoAgentBench: time per task
| Model | Creator | GeoAgentBench: time per task | Coverage (attempts) |
|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 5.4 min | 50 of 150 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 5.5 min | 50 of 150 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 5.2 min | 49 of 150 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 11.2 min | 50 of 150 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 5.9 min | 50 of 150 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 3.0 min | 50 of 150 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 18.5 min | 48 of 150 planned attempts |
Background
GeoAgentBench gives an agent a geospatial question and the data to answer it, and expects the result files and a map. Its tasks are multi-step workflows: reprojecting, overlaying, buffering, raster algebra and network analysis, often ending in a figure.
We rerun each task's recorded reference toolchain with the source's own toolbox to produce the reference result, then check the agent's outputs against it with stated tolerances. Figures are graded by an AI judge against a private checklist.
Published by geox-lab.
