Earth-Bench benchmark leaderboard
Tool use to reach an answer: 48 tasks from Earth-Bench, run by Axis Spatial with the same agent and limits for every model. One of the 4 benchmarks in the Geospatial Agent Index.
Models7 of 7 models
Score
Earth-Bench score
Average pass@1 over its tasks, 0 to 100 · Higher is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
How to read this chart
What this metric means. For each task, the share of a model's attempts that passed every check; then the average over the tasks. Tasks with no decided attempt yet are left out.
Data table: Earth-Bench score
| Model | Creator | Earth-Bench score | Range while attempts are pending | Coverage (attempts) |
|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 40 | No pending judgements; coverage incomplete | 48 of 144 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 48 | No pending judgements; coverage incomplete | 48 of 144 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 30 | 30 to 31 | 55 of 144 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 48 | No pending judgements; coverage incomplete | 48 of 144 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 46 | No pending judgements; coverage incomplete | 48 of 144 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 62 | No pending judgements; coverage incomplete | 48 of 144 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 27 | No pending judgements; coverage incomplete | 48 of 144 planned attempts |
| Model | Score | Tasks with a decided attempt | Attempts decided | Passed | Time limit reached |
|---|---|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | 40 | 48 of 48 | 48 | 19 | 19 |
| Terminus-2 – DeepSeek V4 Pro (high) | 48 | 48 of 48 | 48 | 23 | 14 |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 30 | 48 of 48 | 55 | 18 | 18 |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 48 | 48 of 48 | 48 | 23 | 5 |
| Terminus-2 – gpt-oss-120b (high) | 46 | 48 of 48 | 48 | 22 | 1 |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 62 | 48 of 48 | 48 | 30 | 3 |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 27 | 48 of 48 | 48 | 13 | 32 |
Token usage
Earth-Bench: token usage per attempt
Average tokens per attempt · Lower is better
- Input (not cached)
- Cached input
- Output, including reasoning
- Partial coverage
How to read this chart
What this metric means. Average tokens per attempt, as reported by the provider: input the model read (split into new and cached input) and output it wrote, which includes its reasoning.
Compare with care. Models use different tokenisers, so the same text is a different number of tokens for different models. Use token counts to see how much work each model did, and cost to compare spending.
Data table: Earth-Bench: token usage per attempt
| Model | Input (not cached) | Cached input | Output, including reasoning | Total | Coverage (attempts) |
|---|---|---|---|---|---|
| Terminus-2 – Kimi K2.7 Code (Reasoning) | 42k | 726k | 28k | 796k | 48 of 144 planned attempts |
| Terminus-2 – DeepSeek V4 Flash (high) | 42k | 501k | 61k | 604k | 48 of 144 planned attempts |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | 390k | No data | 18k | 408k | 48 of 144 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | 256k | 63k | 32k | 351k | 48 of 144 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | 81k | No data | 18k | 98k | 48 of 144 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 65k | 8k | 23k | 96k | 56 of 144 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | 64k | No data | 10k | 73k | 48 of 144 planned attempts |
Cost
Earth-Bench: cost per task
Average cost per attempt (USD) · Lower is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
How to read this chart
What cost is measuring. The average cost of one attempt at a task: tokens reported by the provider multiplied by the serving provider's published list prices, with cached input at its own price. Failed attempts count, so a model that fails slowly and expensively is not flattered.
How to read this chart. Lower is better. Costs are pay-per-token list prices, not any discount or credit a customer might have.
Data table: Earth-Bench: cost per task
| Model | Creator | Earth-Bench: cost per task | Coverage (attempts) |
|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | $0.106 | 48 of 144 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | $0.468 | 48 of 144 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | $0.014 | 56 of 144 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | $0.031 | 48 of 144 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | $0.041 | 48 of 144 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | $0.288 | 48 of 144 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | $0.060 | 48 of 144 planned attempts |
Time
Earth-Bench: time per task
Average agent wall time per attempt · Lower is better
- DeepSeek
- Z.ai
- OpenAI
- Moonshot AI
- Alibaba (Qwen)
- Partial coverage
- Terminus-2
- Partial coverage
How to read this chart
What execution time is measuring. Agent wall time per attempt, from the agent's first model call to its last action: model thinking, waiting for the provider and running commands. Sandbox start-up and grading are not included.
How to read this chart. Lower is better. An attempt that reaches the task's time limit scores 0 and counts at the limit.
Data table: Earth-Bench: time per task
| Model | Creator | Earth-Bench: time per task | Coverage (attempts) |
|---|---|---|---|
| Terminus-2 – DeepSeek V4 Flash (high) | DeepSeek | 18.2 min | 48 of 144 planned attempts |
| Terminus-2 – DeepSeek V4 Pro (high) | DeepSeek | 15.9 min | 48 of 144 planned attempts |
| Terminus-2 – Gemma 4 26B A4B (Reasoning) | 15.2 min | 56 of 144 planned attempts | |
| Terminus-2 – GLM-4.7-Flash (Reasoning) | Z.ai | 10.0 min | 48 of 144 planned attempts |
| Terminus-2 – gpt-oss-120b (high) | OpenAI | 5.8 min | 48 of 144 planned attempts |
| Terminus-2 – Kimi K2.7 Code (Reasoning) | Moonshot AI | 12.5 min | 48 of 144 planned attempts |
| Terminus-2 – Qwen 3.8 27B (xhigh) | Alibaba (Qwen) | 21.9 min | 48 of 144 planned attempts |
Background
Earth-Bench poses Earth observation questions over satellite products (indices, temperatures, land cover) that are answered by computing over the data.
The source benchmark gives the agent ready-made tools; here the agent writes Python instead, a disclosed adaptation. Questions that need the model to look at an RGB image are not included.
Published by OpenDataLab.
