Axis Spatial

GeoBenchX benchmark leaderboard

Multi-step spatial execution, including tasks to reject: 173 tasks from GeoBenchX, run by Axis Spatial with the same agent and limits for every model. One of the 4 benchmarks in the Geospatial Agent Index.

Score

GeoBenchX score

Average pass@1 over its tasks, 0 to 100 · Higher is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
GeoBenchX scoreTerminus-2 – Gemma 4 26B A4B (Reasoning): 55 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 54 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 46 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 45 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 44 (partial coverage); Terminus-2 – gpt-oss-120b (high): 42 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): No data01530456055Terminus-2Gemma 4 26BA4B(Reasoning)54Terminus-2Kimi K2.7Code(Reasoning)46Terminus-2GLM-4.7-Flash(Reasoning)45Terminus-2DeepSeek V4Pro (high)44Terminus-2DeepSeek V4Flash (high)42Terminus-2gpt-oss-120b(high)No dataTerminus-2Qwen 3.8 27B(xhigh)
How to read this chart

What this metric means. For each task, the share of a model's attempts that passed every check; then the average over the tasks. Tasks with no decided attempt yet are left out.

Data table: GeoBenchX score
GeoBenchX score
ModelCreatorGeoBenchX scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek4443 to 4652 of 519 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek45No pending judgements; coverage incomplete78 of 519 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google55No pending judgements; coverage incomplete22 of 519 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai4644 to 4826 of 519 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI42No pending judgements; coverage incomplete132 of 519 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI54No pending judgements; coverage incomplete132 of 519 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)No dataNo data0 of 519 planned attempts
GeoBenchX: attempts per model
ModelScoreTasks with a decided attemptAttempts decidedPassedTime limit reached
Terminus-2 – DeepSeek V4 Flash (high)4452 of 17352230
Terminus-2 – DeepSeek V4 Pro (high)4578 of 17378353
Terminus-2 – Gemma 4 26B A4B (Reasoning)5522 of 17322120
Terminus-2 – GLM-4.7-Flash (Reasoning)4626 of 17326121
Terminus-2 – gpt-oss-120b (high)42132 of 173132561
Terminus-2 – Kimi K2.7 Code (Reasoning)54132 of 173132712
Terminus-2 – Qwen 3.8 27B (xhigh)No data0 of 173000

Token usage

GeoBenchX: token usage per attempt

Average tokens per attempt · Lower is better

  • Input (not cached)
  • Cached input
  • Output, including reasoning
  • Partial coverage
GeoBenchX: token usage per attemptTerminus-2 – DeepSeek V4 Flash (high): Input (not cached) 52k, Cached input 233k, Output, including reasoning 28k; Terminus-2 – GLM-4.7-Flash (Reasoning): Input (not cached) 300k, Cached input No data, Output, including reasoning 9k; Terminus-2 – Kimi K2.7 Code (Reasoning): Input (not cached) 23k, Cached input 225k, Output, including reasoning 12k; Terminus-2 – DeepSeek V4 Pro (high): Input (not cached) 164k, Cached input 46k, Output, including reasoning 17k; Terminus-2 – Gemma 4 26B A4B (Reasoning): Input (not cached) 70k, Cached input 23k, Output, including reasoning 9k; Terminus-2 – gpt-oss-120b (high): Input (not cached) 42k, Cached input No data, Output, including reasoning 10k; Terminus-2 – Qwen 3.8 27B (xhigh): No data0100k200k300k400k312kTerminus-2DeepSeek V4Flash (high)310kTerminus-2GLM-4.7-Flash(Reasoning)260kTerminus-2Kimi K2.7Code(Reasoning)226kTerminus-2DeepSeek V4Pro (high)102kTerminus-2Gemma 4 26BA4B(Reasoning)52kTerminus-2gpt-oss-120b(high)No dataTerminus-2Qwen 3.8 27B(xhigh)
How to read this chart

What this metric means. Average tokens per attempt, as reported by the provider: input the model read (split into new and cached input) and output it wrote, which includes its reasoning.

Compare with care. Models use different tokenisers, so the same text is a different number of tokens for different models. Use token counts to see how much work each model did, and cost to compare spending.

Data table: GeoBenchX: token usage per attempt
GeoBenchX: token usage per attempt
ModelInput (not cached)Cached inputOutput, including reasoningTotalCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)52k233k28k312k54 of 519 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)300kNo data9k310k27 of 519 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)23k225k12k260k132 of 519 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)164k46k17k226k78 of 519 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)70k23k9k102k22 of 519 planned attempts
Terminus-2 – gpt-oss-120b (high)42kNo data10k52k132 of 519 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)No dataNo dataNo dataNo data0 of 519 planned attempts

Cost

GeoBenchX: cost per task

Average cost per attempt (USD) · Lower is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
GeoBenchX: cost per taskTerminus-2 – Gemma 4 26B A4B (Reasoning): $0.012 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): $0.022 (partial coverage); Terminus-2 – gpt-oss-120b (high): $0.022 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): $0.063 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): $0.112 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): $0.284 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): No data$0$0.075$0.150$0.225$0.300$0.012Terminus-2Gemma 4 26BA4B(Reasoning)$0.022Terminus-2GLM-4.7-Flash(Reasoning)$0.022Terminus-2gpt-oss-120b(high)$0.063Terminus-2DeepSeek V4Flash (high)$0.112Terminus-2Kimi K2.7Code(Reasoning)$0.284Terminus-2DeepSeek V4Pro (high)No dataTerminus-2Qwen 3.8 27B(xhigh)
How to read this chart

What cost is measuring. The average cost of one attempt at a task: tokens reported by the provider multiplied by the serving provider's published list prices, with cached input at its own price. Failed attempts count, so a model that fails slowly and expensively is not flattered.

How to read this chart. Lower is better. Costs are pay-per-token list prices, not any discount or credit a customer might have.

Data table: GeoBenchX: cost per task
GeoBenchX: cost per task
ModelCreatorGeoBenchX: cost per taskCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek$0.06354 of 519 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek$0.28478 of 519 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google$0.01222 of 519 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai$0.02227 of 519 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI$0.022132 of 519 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI$0.112132 of 519 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)No data0 of 519 planned attempts

Time

GeoBenchX: time per task

Average agent wall time per attempt · Lower is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
GeoBenchX: time per taskTerminus-2 – Gemma 4 26B A4B (Reasoning): 3.4 min (partial coverage); Terminus-2 – gpt-oss-120b (high): 3.7 min (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 6.3 min (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 6.7 min (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 7.3 min (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 8.8 min (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): No data0m2m5m8m10m3.4 minTerminus-2Gemma 4 26BA4B(Reasoning)3.7 minTerminus-2gpt-oss-120b(high)6.3 minTerminus-2Kimi K2.7Code(Reasoning)6.7 minTerminus-2GLM-4.7-Flash(Reasoning)7.3 minTerminus-2DeepSeek V4Pro (high)8.8 minTerminus-2DeepSeek V4Flash (high)No dataTerminus-2Qwen 3.8 27B(xhigh)
How to read this chart

What execution time is measuring. Agent wall time per attempt, from the agent's first model call to its last action: model thinking, waiting for the provider and running commands. Sandbox start-up and grading are not included.

How to read this chart. Lower is better. An attempt that reaches the task's time limit scores 0 and counts at the limit.

Data table: GeoBenchX: time per task
GeoBenchX: time per task
ModelCreatorGeoBenchX: time per taskCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek8.8 min54 of 519 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek7.3 min78 of 519 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google3.4 min22 of 519 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai6.7 min27 of 519 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI3.7 min132 of 519 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI6.3 min132 of 519 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)No data0 of 519 planned attempts

Background

GeoBenchX asks for maps, numbers and lists from country-level and city-level geospatial data. Some of its tasks cannot be done with the data provided; for those, the right answer is to say so.

A rejection is checked by code: rejecting a task that cannot be done passes, and so does solving one that can. Every task asks for the same output files, so the instruction does not reveal which kind it is.

Published by Solirinai.

Related links

Explore evaluations