GeoAnalystBench benchmark leaderboard
Spatial workflow and code generation: 19 tasks from GeoAnalystBench, run by Axis Spatial with the same agent and limits for every model. One of the 4 benchmarks in the Geospatial Agent Index.
Models7 of 7 models
Score
GeoAnalystBench score
Average pass@1 over its tasks, 0 to 100 · Higher is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
How to read this chart
What this metric means. For each task, the share of a model's attempts that passed every check; then the average over the tasks. Tasks with no decided attempt yet are left out.
Data table: GeoAnalystBench score
| Model | Creator | GeoAnalystBench score | Range while attempts are pending | Coverage (attempts) |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 100 | No pending judgements; coverage incomplete | 19 of 57 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 100 | No pending judgements; coverage incomplete | 19 of 57 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 68 | No pending judgements; coverage incomplete | 19 of 57 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 42 | No pending judgements; coverage incomplete | 19 of 57 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 53 | No pending judgements; coverage incomplete | 19 of 57 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 84 | No pending judgements; coverage incomplete | 19 of 57 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 100 | No pending judgements; coverage incomplete | 5 of 57 planned attempts |
| Model | Score | Tasks with a decided attempt | Attempts decided | Passed | Time limit reached |
|---|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | 100 | 19 of 19 | 19 | 19 | 0 |
| Terminus-2 – DeepSeek V4 Pro (high) | 100 | 19 of 19 | 19 | 19 | 0 |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 68 | 19 of 19 | 19 | 13 | 0 |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 42 | 19 of 19 | 19 | 8 | 2 |
| Terminus-2 – gpt-oss-120b (high) | 53 | 19 of 19 | 19 | 10 | 0 |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 84 | 19 of 19 | 19 | 16 | 0 |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 100 | 5 of 19 | 5 | 5 | 0 |
Token usage
GeoAnalystBench: token usage per attempt
Average tokens per attempt · Lower is better
- Input (not cached)
- Cached input
- Output, including reasoning
- Partial coverage
How to read this chart
What this metric means. Average tokens per attempt, as reported by the provider: input the model read (split into new and cached input) and output it wrote, which includes its reasoning.
Compare with care. Models use different tokenisers, so the same text is a different number of tokens for different models. Use token counts to see how much work each model did, and cost to compare spending.
Data table: GeoAnalystBench: token usage per attempt
| Model | Input (not cached) | Cached input | Output, including reasoning | Total | Coverage (attempts) |
|---|---|---|---|---|---|
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 227k | No data | 9k | 237k | 19 of 57 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 102k | No data | 9k | 111k | 5 of 57 planned attempts |
| Terminus-2 – DeepSeek V4 Flash (high) | 12k | 57k | 13k | 82k | 19 of 57 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 8k | 54k | 6k | 67k | 19 of 57 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | 44k | 8k | 7k | 59k | 19 of 57 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 31k | 10k | 10k | 51k | 19 of 57 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | 30k | No data | 10k | 41k | 19 of 57 planned attempts |
Cost
GeoAnalystBench: cost per task
Average cost per attempt (USD) · Lower is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
How to read this chart
What cost is measuring. The average cost of one attempt at a task: tokens reported by the provider multiplied by the serving provider's published list prices, with cached input at its own price. Failed attempts count, so a model that fails slowly and expensively is not flattered.
How to read this chart. Lower is better. Costs are pay-per-token list prices, not any discount or credit a customer might have.
Data table: GeoAnalystBench: cost per task
| Model | Creator | GeoAnalystBench: cost per task | Coverage (attempts) |
|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | $0.023 | 19 of 57 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | $0.087 | 19 of 57 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | $0.0072 | 19 of 57 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | $0.017 | 19 of 57 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | $0.018 | 19 of 57 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | $0.040 | 19 of 57 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | $0.075 | 5 of 57 planned attempts |
Time
GeoAnalystBench: time per task
Average agent wall time per attempt · Lower is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
How to read this chart
What execution time is measuring. Agent wall time per attempt, from the agent's first model call to its last action: model thinking, waiting for the provider and running commands. Sandbox start-up and grading are not included.
How to read this chart. Lower is better. An attempt that reaches the task's time limit scores 0 and counts at the limit.
Data table: GeoAnalystBench: time per task
| Model | Creator | GeoAnalystBench: time per task | Coverage (attempts) |
|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 4.6 min | 19 of 57 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 3.3 min | 19 of 57 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 4.2 min | 19 of 57 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 7.3 min | 19 of 57 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 3.6 min | 19 of 57 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 4.1 min | 19 of 57 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 10.7 min | 5 of 57 planned attempts |
Background
GeoAnalystBench collects GIS analyst workflows (suitability analysis, interpolation, hot spots, network and overlay work). It was built to test workflow and code generation: the source benchmark asks a model to write the workflow and the code for each task.
Here the agent writes and runs the workflow itself, and is graded on the resulting data and maps, so these tasks count as multi-step analysis, a disclosed adaptation. Tasks that need ArcPy, a proprietary library, are not included; the others run with open tools in a sandbox without internet access.
Published by GeoDS Lab.
