Axis Spatial

Earth-Bench benchmark leaderboard

Tool use to reach an answer: 48 tasks from Earth-Bench, run by Axis Spatial with the same agent and limits for every model. One of the 4 benchmarks in the Geospatial Agent Index.

Score

Earth-Bench score

Average pass@1 over its tasks, 0 to 100 · Higher is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
Earth-Bench scoreTerminus-2 – Kimi K2.7 Code (Reasoning): 62 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 48 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 48 (partial coverage); Terminus-2 – gpt-oss-120b (high): 46 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 40 (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 30 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 27 (partial coverage)02040608062Terminus-2Kimi K2.7Code(Reasoning)48Terminus-2DeepSeek V4Pro (high)48Terminus-2GLM-4.7-Flash(Reasoning)46Terminus-2gpt-oss-120b(high)40Terminus-2DeepSeek V4Flash (high)30Terminus-2Gemma 4 26BA4B(Reasoning)27Terminus-2Qwen 3.8 27B(xhigh)
How to read this chart

What this metric means. For each task, the share of a model's attempts that passed every check; then the average over the tasks. Tasks with no decided attempt yet are left out.

Data table: Earth-Bench score
Earth-Bench score
ModelCreatorEarth-Bench scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek40No pending judgements; coverage incomplete48 of 144 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek48No pending judgements; coverage incomplete48 of 144 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google3030 to 3155 of 144 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai48No pending judgements; coverage incomplete48 of 144 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI46No pending judgements; coverage incomplete48 of 144 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI62No pending judgements; coverage incomplete48 of 144 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)27No pending judgements; coverage incomplete48 of 144 planned attempts
Earth-Bench: attempts per model
ModelScoreTasks with a decided attemptAttempts decidedPassedTime limit reached
Terminus-2 – DeepSeek V4 Flash (high)4048 of 48481919
Terminus-2 – DeepSeek V4 Pro (high)4848 of 48482314
Terminus-2 – Gemma 4 26B A4B (Reasoning)3048 of 48551818
Terminus-2 – GLM-4.7-Flash (Reasoning)4848 of 4848235
Terminus-2 – gpt-oss-120b (high)4648 of 4848221
Terminus-2 – Kimi K2.7 Code (Reasoning)6248 of 4848303
Terminus-2 – Qwen 3.8 27B (xhigh)2748 of 48481332

Token usage

Earth-Bench: token usage per attempt

Average tokens per attempt · Lower is better

  • Input (not cached)
  • Cached input
  • Output, including reasoning
  • Partial coverage
Earth-Bench: token usage per attemptTerminus-2 – Kimi K2.7 Code (Reasoning): Input (not cached) 42k, Cached input 726k, Output, including reasoning 28k; Terminus-2 – DeepSeek V4 Flash (high): Input (not cached) 42k, Cached input 501k, Output, including reasoning 61k; Terminus-2 – GLM-4.7-Flash (Reasoning): Input (not cached) 390k, Cached input No data, Output, including reasoning 18k; Terminus-2 – DeepSeek V4 Pro (high): Input (not cached) 256k, Cached input 63k, Output, including reasoning 32k; Terminus-2 – gpt-oss-120b (high): Input (not cached) 81k, Cached input No data, Output, including reasoning 18k; Terminus-2 – Gemma 4 26B A4B (Reasoning): Input (not cached) 65k, Cached input 8k, Output, including reasoning 23k; Terminus-2 – Qwen 3.8 27B (xhigh): Input (not cached) 64k, Cached input No data, Output, including reasoning 10k0200k400k600k800k796kTerminus-2Kimi K2.7Code(Reasoning)604kTerminus-2DeepSeek V4Flash (high)408kTerminus-2GLM-4.7-Flash(Reasoning)351kTerminus-2DeepSeek V4Pro (high)98kTerminus-2gpt-oss-120b(high)96kTerminus-2Gemma 4 26BA4B(Reasoning)73kTerminus-2Qwen 3.8 27B(xhigh)
How to read this chart

What this metric means. Average tokens per attempt, as reported by the provider: input the model read (split into new and cached input) and output it wrote, which includes its reasoning.

Compare with care. Models use different tokenisers, so the same text is a different number of tokens for different models. Use token counts to see how much work each model did, and cost to compare spending.

Data table: Earth-Bench: token usage per attempt
Earth-Bench: token usage per attempt
ModelInput (not cached)Cached inputOutput, including reasoningTotalCoverage (attempts)
Terminus-2 – Kimi K2.7 Code (Reasoning)42k726k28k796k48 of 144 planned attempts
Terminus-2 – DeepSeek V4 Flash (high)42k501k61k604k48 of 144 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)390kNo data18k408k48 of 144 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)256k63k32k351k48 of 144 planned attempts
Terminus-2 – gpt-oss-120b (high)81kNo data18k98k48 of 144 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)65k8k23k96k56 of 144 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)64kNo data10k73k48 of 144 planned attempts

Cost

Earth-Bench: cost per task

Average cost per attempt (USD) · Lower is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
Earth-Bench: cost per taskTerminus-2 – Gemma 4 26B A4B (Reasoning): $0.014 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): $0.031 (partial coverage); Terminus-2 – gpt-oss-120b (high): $0.041 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): $0.060 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): $0.106 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): $0.288 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): $0.468 (partial coverage)$0$0.125$0.250$0.375$0.500$0.014Terminus-2Gemma 4 26BA4B(Reasoning)$0.031Terminus-2GLM-4.7-Flash(Reasoning)$0.041Terminus-2gpt-oss-120b(high)$0.060Terminus-2Qwen 3.8 27B(xhigh)$0.106Terminus-2DeepSeek V4Flash (high)$0.288Terminus-2Kimi K2.7Code(Reasoning)$0.468Terminus-2DeepSeek V4Pro (high)
How to read this chart

What cost is measuring. The average cost of one attempt at a task: tokens reported by the provider multiplied by the serving provider's published list prices, with cached input at its own price. Failed attempts count, so a model that fails slowly and expensively is not flattered.

How to read this chart. Lower is better. Costs are pay-per-token list prices, not any discount or credit a customer might have.

Data table: Earth-Bench: cost per task
Earth-Bench: cost per task
ModelCreatorEarth-Bench: cost per taskCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek$0.10648 of 144 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek$0.46848 of 144 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google$0.01456 of 144 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai$0.03148 of 144 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI$0.04148 of 144 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI$0.28848 of 144 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)$0.06048 of 144 planned attempts

Time

Earth-Bench: time per task

Average agent wall time per attempt · Lower is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
Earth-Bench: time per taskTerminus-2 – gpt-oss-120b (high): 5.8 min (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 10.0 min (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 12.5 min (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 15.2 min (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 15.9 min (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 18.2 min (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 21.9 min (partial coverage)0m6m12m19m25m5.8 minTerminus-2gpt-oss-120b(high)10.0 minTerminus-2GLM-4.7-Flash(Reasoning)12.5 minTerminus-2Kimi K2.7Code(Reasoning)15.2 minTerminus-2Gemma 4 26BA4B(Reasoning)15.9 minTerminus-2DeepSeek V4Pro (high)18.2 minTerminus-2DeepSeek V4Flash (high)21.9 minTerminus-2Qwen 3.8 27B(xhigh)
How to read this chart

What execution time is measuring. Agent wall time per attempt, from the agent's first model call to its last action: model thinking, waiting for the provider and running commands. Sandbox start-up and grading are not included.

How to read this chart. Lower is better. An attempt that reaches the task's time limit scores 0 and counts at the limit.

Data table: Earth-Bench: time per task
Earth-Bench: time per task
ModelCreatorEarth-Bench: time per taskCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek18.2 min48 of 144 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek15.9 min48 of 144 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google15.2 min56 of 144 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai10.0 min48 of 144 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI5.8 min48 of 144 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI12.5 min48 of 144 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)21.9 min48 of 144 planned attempts

Background

Earth-Bench poses Earth observation questions over satellite products (indices, temperatures, land cover) that are answered by computing over the data.

The source benchmark gives the agent ready-made tools; here the agent writes Python instead, a disclosed adaptation. Questions that need the model to look at an RGB image are not included.

Published by OpenDataLab.

Related links

Explore evaluations