GeoBenchX benchmark leaderboard
Multi-step spatial execution, including tasks to reject: 173 tasks from GeoBenchX, run by Axis Spatial with the same agent and limits for every model. One of the 4 benchmarks in the Geospatial Agent Index.
Models7 of 7 models
Score
GeoBenchX score
Average pass@1 over its tasks, 0 to 100 · Higher is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
How to read this chart
What this metric means. For each task, the share of a model's attempts that passed every check; then the average over the tasks. Tasks with no decided attempt yet are left out.
Data table: GeoBenchX score
| Model | Creator | GeoBenchX score | Range while attempts are pending | Coverage (attempts) |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 44 | 43 to 46 | 52 of 519 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 45 | No pending judgements; coverage incomplete | 78 of 519 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 55 | No pending judgements; coverage incomplete | 22 of 519 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 46 | 44 to 48 | 26 of 519 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 42 | No pending judgements; coverage incomplete | 132 of 519 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 54 | No pending judgements; coverage incomplete | 132 of 519 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | No data | No data | 0 of 519 planned attempts |
| Model | Score | Tasks with a decided attempt | Attempts decided | Passed | Time limit reached |
|---|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | 44 | 52 of 173 | 52 | 23 | 0 |
| Terminus-2 – DeepSeek V4 Pro (high) | 45 | 78 of 173 | 78 | 35 | 3 |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 55 | 22 of 173 | 22 | 12 | 0 |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 46 | 26 of 173 | 26 | 12 | 1 |
| Terminus-2 – gpt-oss-120b (high) | 42 | 132 of 173 | 132 | 56 | 1 |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 54 | 132 of 173 | 132 | 71 | 2 |
| Terminus-2 – Qwen 3.8 27B (xhigh) | No data | 0 of 173 | 0 | 0 | 0 |
Token usage
GeoBenchX: token usage per attempt
Average tokens per attempt · Lower is better
- Input (not cached)
- Cached input
- Output, including reasoning
- Partial coverage
How to read this chart
What this metric means. Average tokens per attempt, as reported by the provider: input the model read (split into new and cached input) and output it wrote, which includes its reasoning.
Compare with care. Models use different tokenisers, so the same text is a different number of tokens for different models. Use token counts to see how much work each model did, and cost to compare spending.
Data table: GeoBenchX: token usage per attempt
| Model | Input (not cached) | Cached input | Output, including reasoning | Total | Coverage (attempts) |
|---|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | 52k | 233k | 28k | 312k | 54 of 519 planned attempts |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 300k | No data | 9k | 310k | 27 of 519 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 23k | 225k | 12k | 260k | 132 of 519 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | 164k | 46k | 17k | 226k | 78 of 519 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 70k | 23k | 9k | 102k | 22 of 519 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | 42k | No data | 10k | 52k | 132 of 519 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | No data | No data | No data | No data | 0 of 519 planned attempts |
Cost
GeoBenchX: cost per task
Average cost per attempt (USD) · Lower is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
How to read this chart
What cost is measuring. The average cost of one attempt at a task: tokens reported by the provider multiplied by the serving provider's published list prices, with cached input at its own price. Failed attempts count, so a model that fails slowly and expensively is not flattered.
How to read this chart. Lower is better. Costs are pay-per-token list prices, not any discount or credit a customer might have.
Data table: GeoBenchX: cost per task
| Model | Creator | GeoBenchX: cost per task | Coverage (attempts) |
|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | $0.063 | 54 of 519 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | $0.284 | 78 of 519 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | $0.012 | 22 of 519 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | $0.022 | 27 of 519 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | $0.022 | 132 of 519 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | $0.112 | 132 of 519 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | No data | 0 of 519 planned attempts |
Time
GeoBenchX: time per task
Average agent wall time per attempt · Lower is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
How to read this chart
What execution time is measuring. Agent wall time per attempt, from the agent's first model call to its last action: model thinking, waiting for the provider and running commands. Sandbox start-up and grading are not included.
How to read this chart. Lower is better. An attempt that reaches the task's time limit scores 0 and counts at the limit.
Data table: GeoBenchX: time per task
| Model | Creator | GeoBenchX: time per task | Coverage (attempts) |
|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 8.8 min | 54 of 519 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 7.3 min | 78 of 519 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 3.4 min | 22 of 519 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 6.7 min | 27 of 519 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 3.7 min | 132 of 519 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 6.3 min | 132 of 519 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | No data | 0 of 519 planned attempts |
Background
GeoBenchX asks for maps, numbers and lists from country-level and city-level geospatial data. Some of its tasks cannot be done with the data provided; for those, the right answer is to say so.
A rejection is checked by code: rejecting a task that cannot be done passes, and so does solving one that can. Every task asks for the same output files, so the instruction does not reveal which kind it is.
Published by Solirinai.
