Axis Spatial

GeoAnalystBench benchmark leaderboard

Spatial workflow and code generation: 19 tasks from GeoAnalystBench, run by Axis Spatial with the same agent and limits for every model. One of the 4 benchmarks in the Geospatial Agent Index.

Score

GeoAnalystBench score

Average pass@1 over its tasks, 0 to 100 · Higher is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
GeoAnalystBench scoreTerminus-2 – DeepSeek V4 Flash (high): 100 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 100 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 100 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 84 (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 68 (partial coverage); Terminus-2 – gpt-oss-120b (high): 53 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 42 (partial coverage)0255075100100Terminus-2DeepSeek V4Flash (high)100Terminus-2DeepSeek V4Pro (high)100Terminus-2Qwen 3.8 27B(xhigh)84Terminus-2Kimi K2.7Code(Reasoning)68Terminus-2Gemma 4 26BA4B(Reasoning)53Terminus-2gpt-oss-120b(high)42Terminus-2GLM-4.7-Flash(Reasoning)
How to read this chart

What this metric means. For each task, the share of a model's attempts that passed every check; then the average over the tasks. Tasks with no decided attempt yet are left out.

Data table: GeoAnalystBench score
GeoAnalystBench score
ModelCreatorGeoAnalystBench scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek100No pending judgements; coverage incomplete19 of 57 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek100No pending judgements; coverage incomplete19 of 57 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google68No pending judgements; coverage incomplete19 of 57 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai42No pending judgements; coverage incomplete19 of 57 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI53No pending judgements; coverage incomplete19 of 57 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI84No pending judgements; coverage incomplete19 of 57 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)100No pending judgements; coverage incomplete5 of 57 planned attempts
GeoAnalystBench: attempts per model
ModelScoreTasks with a decided attemptAttempts decidedPassedTime limit reached
Terminus-2 – DeepSeek V4 Flash (high)10019 of 1919190
Terminus-2 – DeepSeek V4 Pro (high)10019 of 1919190
Terminus-2 – Gemma 4 26B A4B (Reasoning)6819 of 1919130
Terminus-2 – GLM-4.7-Flash (Reasoning)4219 of 191982
Terminus-2 – gpt-oss-120b (high)5319 of 1919100
Terminus-2 – Kimi K2.7 Code (Reasoning)8419 of 1919160
Terminus-2 – Qwen 3.8 27B (xhigh)1005 of 19550

Token usage

GeoAnalystBench: token usage per attempt

Average tokens per attempt · Lower is better

  • Input (not cached)
  • Cached input
  • Output, including reasoning
  • Partial coverage
GeoAnalystBench: token usage per attemptTerminus-2 – GLM-4.7-Flash (Reasoning): Input (not cached) 227k, Cached input No data, Output, including reasoning 9k; Terminus-2 – Qwen 3.8 27B (xhigh): Input (not cached) 102k, Cached input No data, Output, including reasoning 9k; Terminus-2 – DeepSeek V4 Flash (high): Input (not cached) 12k, Cached input 57k, Output, including reasoning 13k; Terminus-2 – Kimi K2.7 Code (Reasoning): Input (not cached) 8k, Cached input 54k, Output, including reasoning 6k; Terminus-2 – DeepSeek V4 Pro (high): Input (not cached) 44k, Cached input 8k, Output, including reasoning 7k; Terminus-2 – Gemma 4 26B A4B (Reasoning): Input (not cached) 31k, Cached input 10k, Output, including reasoning 10k; Terminus-2 – gpt-oss-120b (high): Input (not cached) 30k, Cached input No data, Output, including reasoning 10k062k125k188k250k237kTerminus-2GLM-4.7-Flash(Reasoning)111kTerminus-2Qwen 3.8 27B(xhigh)82kTerminus-2DeepSeek V4Flash (high)67kTerminus-2Kimi K2.7Code(Reasoning)59kTerminus-2DeepSeek V4Pro (high)51kTerminus-2Gemma 4 26BA4B(Reasoning)41kTerminus-2gpt-oss-120b(high)
How to read this chart

What this metric means. Average tokens per attempt, as reported by the provider: input the model read (split into new and cached input) and output it wrote, which includes its reasoning.

Compare with care. Models use different tokenisers, so the same text is a different number of tokens for different models. Use token counts to see how much work each model did, and cost to compare spending.

Data table: GeoAnalystBench: token usage per attempt
GeoAnalystBench: token usage per attempt
ModelInput (not cached)Cached inputOutput, including reasoningTotalCoverage (attempts)
Terminus-2 – GLM-4.7-Flash (Reasoning)227kNo data9k237k19 of 57 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)102kNo data9k111k5 of 57 planned attempts
Terminus-2 – DeepSeek V4 Flash (high)12k57k13k82k19 of 57 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)8k54k6k67k19 of 57 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)44k8k7k59k19 of 57 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)31k10k10k51k19 of 57 planned attempts
Terminus-2 – gpt-oss-120b (high)30kNo data10k41k19 of 57 planned attempts

Cost

GeoAnalystBench: cost per task

Average cost per attempt (USD) · Lower is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
GeoAnalystBench: cost per taskTerminus-2 – Gemma 4 26B A4B (Reasoning): $0.0072 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): $0.017 (partial coverage); Terminus-2 – gpt-oss-120b (high): $0.018 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): $0.023 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): $0.040 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): $0.075 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): $0.087 (partial coverage)$0$0.025$0.050$0.075$0.100$0.0072Terminus-2Gemma 4 26BA4B(Reasoning)$0.017Terminus-2GLM-4.7-Flash(Reasoning)$0.018Terminus-2gpt-oss-120b(high)$0.023Terminus-2DeepSeek V4Flash (high)$0.040Terminus-2Kimi K2.7Code(Reasoning)$0.075Terminus-2Qwen 3.8 27B(xhigh)$0.087Terminus-2DeepSeek V4Pro (high)
How to read this chart

What cost is measuring. The average cost of one attempt at a task: tokens reported by the provider multiplied by the serving provider's published list prices, with cached input at its own price. Failed attempts count, so a model that fails slowly and expensively is not flattered.

How to read this chart. Lower is better. Costs are pay-per-token list prices, not any discount or credit a customer might have.

Data table: GeoAnalystBench: cost per task
GeoAnalystBench: cost per task
ModelCreatorGeoAnalystBench: cost per taskCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek$0.02319 of 57 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek$0.08719 of 57 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google$0.007219 of 57 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai$0.01719 of 57 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI$0.01819 of 57 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI$0.04019 of 57 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)$0.0755 of 57 planned attempts

Time

GeoAnalystBench: time per task

Average agent wall time per attempt · Lower is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
GeoAnalystBench: time per taskTerminus-2 – DeepSeek V4 Pro (high): 3.3 min (partial coverage); Terminus-2 – gpt-oss-120b (high): 3.6 min (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 4.1 min (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 4.2 min (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 4.6 min (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 7.3 min (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 10.7 min (partial coverage)0m3m7m10m13m3.3 minTerminus-2DeepSeek V4Pro (high)3.6 minTerminus-2gpt-oss-120b(high)4.1 minTerminus-2Kimi K2.7Code(Reasoning)4.2 minTerminus-2Gemma 4 26BA4B(Reasoning)4.6 minTerminus-2DeepSeek V4Flash (high)7.3 minTerminus-2GLM-4.7-Flash(Reasoning)10.7 minTerminus-2Qwen 3.8 27B(xhigh)
How to read this chart

What execution time is measuring. Agent wall time per attempt, from the agent's first model call to its last action: model thinking, waiting for the provider and running commands. Sandbox start-up and grading are not included.

How to read this chart. Lower is better. An attempt that reaches the task's time limit scores 0 and counts at the limit.

Data table: GeoAnalystBench: time per task
GeoAnalystBench: time per task
ModelCreatorGeoAnalystBench: time per taskCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek4.6 min19 of 57 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek3.3 min19 of 57 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google4.2 min19 of 57 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai7.3 min19 of 57 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI3.6 min19 of 57 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI4.1 min19 of 57 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)10.7 min5 of 57 planned attempts

Background

GeoAnalystBench collects GIS analyst workflows (suitability analysis, interpolation, hot spots, network and overlay work). It was built to test workflow and code generation: the source benchmark asks a model to write the workflow and the code for each task.

Here the agent writes and runs the workflow itself, and is graded on the resulting data and maps, so these tasks count as multi-step analysis, a disclosed adaptation. Tasks that need ArcPy, a proprietary library, are not included; the others run with open tools in a sandbox without internet access.

Published by GeoDS Lab.

Related links

Explore evaluations