Axis Spatial

GeoAgentBench benchmark leaderboard

Multi-step spatial execution: 50 tasks from GeoAgentBench, run by Axis Spatial with the same agent and limits for every model. One of the 4 benchmarks in the Geospatial Agent Index.

Score

GeoAgentBench score

Average pass@1 over its tasks, 0 to 100 · Higher is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
GeoAgentBench scoreTerminus-2 – DeepSeek V4 Pro (high): 98 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 96 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 90 (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 86 (partial coverage); Terminus-2 – gpt-oss-120b (high): 70 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 60 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 36 (partial coverage)025507510098Terminus-2DeepSeek V4Pro (high)96Terminus-2DeepSeek V4Flash (high)90Terminus-2Kimi K2.7Code(Reasoning)86Terminus-2Gemma 4 26BA4B(Reasoning)70Terminus-2gpt-oss-120b(high)60Terminus-2Qwen 3.8 27B(xhigh)36Terminus-2GLM-4.7-Flash(Reasoning)
How to read this chart

What this metric means. For each task, the share of a model's attempts that passed every check; then the average over the tasks. Tasks with no decided attempt yet are left out.

Data table: GeoAgentBench score
GeoAgentBench score
ModelCreatorGeoAgentBench scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek96No pending judgements; coverage incomplete50 of 150 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek98No pending judgements; coverage incomplete50 of 150 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google86No pending judgements; coverage incomplete49 of 150 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai36No pending judgements; coverage incomplete50 of 150 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI70No pending judgements; coverage incomplete50 of 150 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI90No pending judgements; coverage incomplete50 of 150 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)60No pending judgements; coverage incomplete48 of 150 planned attempts
GeoAgentBench: attempts per model
ModelScoreTasks with a decided attemptAttempts decidedPassedTime limit reached
Terminus-2 – DeepSeek V4 Flash (high)9650 of 5050480
Terminus-2 – DeepSeek V4 Pro (high)9850 of 5050490
Terminus-2 – Gemma 4 26B A4B (Reasoning)8649 of 5049421
Terminus-2 – GLM-4.7-Flash (Reasoning)3650 of 50501810
Terminus-2 – gpt-oss-120b (high)7050 of 5050352
Terminus-2 – Kimi K2.7 Code (Reasoning)9050 of 5050450
Terminus-2 – Qwen 3.8 27B (xhigh)6048 of 50482919

Token usage

GeoAgentBench: token usage per attempt

Average tokens per attempt · Lower is better

  • Input (not cached)
  • Cached input
  • Output, including reasoning
  • Partial coverage
GeoAgentBench: token usage per attemptTerminus-2 – GLM-4.7-Flash (Reasoning): Input (not cached) 293k, Cached input No data, Output, including reasoning 14k; Terminus-2 – Qwen 3.8 27B (xhigh): Input (not cached) 97k, Cached input No data, Output, including reasoning 8k; Terminus-2 – Gemma 4 26B A4B (Reasoning): Input (not cached) 52k, Cached input 24k, Output, including reasoning 13k; Terminus-2 – DeepSeek V4 Flash (high): Input (not cached) 11k, Cached input 50k, Output, including reasoning 16k; Terminus-2 – Kimi K2.7 Code (Reasoning): Input (not cached) 7k, Cached input 50k, Output, including reasoning 6k; Terminus-2 – DeepSeek V4 Pro (high): Input (not cached) 40k, Cached input 6k, Output, including reasoning 10k; Terminus-2 – gpt-oss-120b (high): Input (not cached) 35k, Cached input No data, Output, including reasoning 14k0100k200k300k400k307kTerminus-2GLM-4.7-Flash(Reasoning)105kTerminus-2Qwen 3.8 27B(xhigh)89kTerminus-2Gemma 4 26BA4B(Reasoning)77kTerminus-2DeepSeek V4Flash (high)63kTerminus-2Kimi K2.7Code(Reasoning)56kTerminus-2DeepSeek V4Pro (high)49kTerminus-2gpt-oss-120b(high)
How to read this chart

What this metric means. Average tokens per attempt, as reported by the provider: input the model read (split into new and cached input) and output it wrote, which includes its reasoning.

Compare with care. Models use different tokenisers, so the same text is a different number of tokens for different models. Use token counts to see how much work each model did, and cost to compare spending.

Data table: GeoAgentBench: token usage per attempt
GeoAgentBench: token usage per attempt
ModelInput (not cached)Cached inputOutput, including reasoningTotalCoverage (attempts)
Terminus-2 – GLM-4.7-Flash (Reasoning)293kNo data14k307k50 of 150 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)97kNo data8k105k48 of 150 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)52k24k13k89k49 of 150 planned attempts
Terminus-2 – DeepSeek V4 Flash (high)11k50k16k77k50 of 150 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)7k50k6k63k50 of 150 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)40k6k10k56k50 of 150 planned attempts
Terminus-2 – gpt-oss-120b (high)35kNo data14k49k50 of 150 planned attempts

Cost

GeoAgentBench: cost per task

Average cost per attempt (USD) · Lower is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
GeoAgentBench: cost per taskTerminus-2 – Gemma 4 26B A4B (Reasoning): $0.012 (partial coverage); Terminus-2 – gpt-oss-120b (high): $0.023 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): $0.023 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): $0.026 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): $0.038 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): $0.070 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): $0.091 (partial coverage)$0$0.025$0.050$0.075$0.100$0.012Terminus-2Gemma 4 26BA4B(Reasoning)$0.023Terminus-2gpt-oss-120b(high)$0.023Terminus-2GLM-4.7-Flash(Reasoning)$0.026Terminus-2DeepSeek V4Flash (high)$0.038Terminus-2Kimi K2.7Code(Reasoning)$0.070Terminus-2Qwen 3.8 27B(xhigh)$0.091Terminus-2DeepSeek V4Pro (high)
How to read this chart

What cost is measuring. The average cost of one attempt at a task: tokens reported by the provider multiplied by the serving provider's published list prices, with cached input at its own price. Failed attempts count, so a model that fails slowly and expensively is not flattered.

How to read this chart. Lower is better. Costs are pay-per-token list prices, not any discount or credit a customer might have.

Data table: GeoAgentBench: cost per task
GeoAgentBench: cost per task
ModelCreatorGeoAgentBench: cost per taskCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek$0.02650 of 150 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek$0.09150 of 150 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google$0.01249 of 150 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai$0.02350 of 150 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI$0.02350 of 150 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI$0.03850 of 150 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)$0.07048 of 150 planned attempts

Time

GeoAgentBench: time per task

Average agent wall time per attempt · Lower is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
GeoAgentBench: time per taskTerminus-2 – Kimi K2.7 Code (Reasoning): 3.0 min (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 5.2 min (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 5.4 min (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 5.5 min (partial coverage); Terminus-2 – gpt-oss-120b (high): 5.9 min (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 11.2 min (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 18.5 min (partial coverage)0m5m10m15m20m3.0 minTerminus-2Kimi K2.7Code(Reasoning)5.2 minTerminus-2Gemma 4 26BA4B(Reasoning)5.4 minTerminus-2DeepSeek V4Flash (high)5.5 minTerminus-2DeepSeek V4Pro (high)5.9 minTerminus-2gpt-oss-120b(high)11.2 minTerminus-2GLM-4.7-Flash(Reasoning)18.5 minTerminus-2Qwen 3.8 27B(xhigh)
How to read this chart

What execution time is measuring. Agent wall time per attempt, from the agent's first model call to its last action: model thinking, waiting for the provider and running commands. Sandbox start-up and grading are not included.

How to read this chart. Lower is better. An attempt that reaches the task's time limit scores 0 and counts at the limit.

Data table: GeoAgentBench: time per task
GeoAgentBench: time per task
ModelCreatorGeoAgentBench: time per taskCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek5.4 min50 of 150 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek5.5 min50 of 150 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google5.2 min49 of 150 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai11.2 min50 of 150 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI5.9 min50 of 150 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI3.0 min50 of 150 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)18.5 min48 of 150 planned attempts

Background

GeoAgentBench gives an agent a geospatial question and the data to answer it, and expects the result files and a map. Its tasks are multi-step workflows: reprojecting, overlaying, buffering, raster algebra and network analysis, often ending in a figure.

We rerun each task's recorded reference toolchain with the source's own toolbox to produce the reference result, then check the agent's outputs against it with stated tolerances. Figures are graded by an AI judge against a private checklist.

Published by geox-lab.

Related links

Explore evaluations