Axis Spatial

Axis Spatial Geospatial Agent Index

The Geospatial Agent Index combines 4 published geospatial benchmarks into one score, weighting each benchmark equally so that a benchmark with many tasks does not outweigh one with few.

Geospatial Agent Index

Composite index of 4 benchmarks, 290 tasks in all:

  • GeoAgentBench
    Multi-step spatial execution, 50 tasks
    by geox-lab
  • GeoBenchX
    Multi-step spatial execution, including tasks to reject, 173 tasks
    by Solirinai
  • Earth-Bench
    Tool use to reach an answer, 48 tasks
    by OpenDataLab
  • GeoAnalystBench
    Spatial workflow and code generation, 19 tasks
    by GeoDS Lab

Each benchmark score is the average pass@1 over its tasks, with three attempts per task and model; the index weights the benchmarks equally. Methodology version 0.4. How we score.

Score

Axis Spatial Geospatial Agent Index

Equal-weight average of the evaluation scores, 0 to 100 · Higher is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
Axis Spatial Geospatial Agent IndexTerminus-2 – DeepSeek V4 Pro (high): 73* (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 73* (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 70* (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 60* (partial coverage); Terminus-2 – gpt-oss-120b (high): 53* (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 43* (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 62* (partial coverage)025507510073*Terminus-2DeepSeek V4Pro (high)73*Terminus-2Kimi K2.7Code(Reasoning)70*Terminus-2DeepSeek V4Flash (high)60*Terminus-2Gemma 4 26BA4B(Reasoning)53*Terminus-2gpt-oss-120b(high)43*Terminus-2GLM-4.7-Flash(Reasoning)62*Terminus-2Qwen 3.8 27B(xhigh)
How to read this chart

What this metric means. The equal-weight average of the 4 evaluation scores. Each evaluation score is the average pass@1 over that benchmark's tasks, from 0 to 100.

Not yet ranked. Entries without a score on every benchmark come after the ranked entries and are not ranked: their index averages only the benchmarks they have.

Partial coverage. A hatched bar, and an asterisk in tables, marks an entry that does not yet have results for every planned attempt. While attempts await review, the table gives the range the score could still reach.

Data table: Axis Spatial Geospatial Agent Index
Axis Spatial Geospatial Agent Index
ModelCreatorGeospatial Agent IndexRange while attempts are pendingBenchmarks coveredCoverage (attempts)
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek73*No pending judgements; coverage incomplete4 of 4195 of 870 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI73*No pending judgements; coverage incomplete4 of 4249 of 870 planned attempts
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek70*70 to 704 of 4169 of 870 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google60*60 to 604 of 4145 of 870 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI53*No pending judgements; coverage incomplete4 of 4249 of 870 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai43*43 to 444 of 4143 of 870 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)62*No pending judgements; coverage incomplete3 of 4101 of 870 planned attempts
Score by benchmark (0 to 100)
ModelGeoAgentBenchGeoBenchXEarth-BenchGeoAnalystBenchGeospatial Agent Index
Terminus-2 – DeepSeek V4 Pro (high)98454810073*
Terminus-2 – Kimi K2.7 Code (Reasoning)9054628473*
Terminus-2 – DeepSeek V4 Flash (high)96444010070*
Terminus-2 – Gemma 4 26B A4B (Reasoning)8655306860*
Terminus-2 – gpt-oss-120b (high)7042465353*
Terminus-2 – GLM-4.7-Flash (Reasoning)3646484243*
Terminus-2 – Qwen 3.8 27B (xhigh)60No data2710062*

Cost and time

Axis Spatial Geospatial Agent Index vs. cost per task

Higher and further left is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
  • Most attractive quadrant
  • Pareto line
Axis Spatial Geospatial Agent Index vs. cost per taskTerminus-2 – DeepSeek V4 Flash (high): Geospatial Agent Index 69.95, cost per task (usd) $0.060; Terminus-2 – DeepSeek V4 Pro (high): Geospatial Agent Index 72.7, cost per task (usd) $0.261; Terminus-2 – Gemma 4 26B A4B (Reasoning): Geospatial Agent Index 59.63, cost per task (usd) $0.012; Terminus-2 – GLM-4.7-Flash (Reasoning): Geospatial Agent Index 43.05, cost per task (usd) $0.025; Terminus-2 – gpt-oss-120b (high): Geospatial Agent Index 52.72, cost per task (usd) $0.026; Terminus-2 – Kimi K2.7 Code (Reasoning): Geospatial Agent Index 72.62, cost per task (usd) $0.126; Terminus-2 – Qwen 3.8 27B (xhigh): Geospatial Agent Index 62.5, cost per task (usd) $0.066Most attractive quadrant0255075100$0.010$0.020$0.050$0.100$0.200$0.500$1.00Cost per task (USD) (log scale)Geospatial Agent IndexDeepSeek V4 Pro (high)Kimi K2.7 Code (Reasoning)DeepSeek V4 Flash (high)Qwen 3.8 27B (xhigh)Gemma 4 26B A4B (Reasoning)gpt-oss-120b (high)GLM-4.7-Flash (Reasoning)
How to read this chart

What cost is measuring. The average cost of one attempt at a task: tokens reported by the provider multiplied by the serving provider's published list prices, with cached input at its own price. Failed attempts count, so a model that fails slowly and expensively is not flattered.

How to read this chart. Lower is better. Costs are pay-per-token list prices, not any discount or credit a customer might have.

Data table: Axis Spatial Geospatial Agent Index vs. cost per task
Axis Spatial Geospatial Agent Index vs. cost per task
ModelGeospatial Agent IndexCost per task (USD)On the Pareto line
Terminus-2 – DeepSeek V4 Flash (high)70*$0.060Yes
Terminus-2 – DeepSeek V4 Pro (high)73*$0.261Yes
Terminus-2 – Gemma 4 26B A4B (Reasoning)60*$0.012Yes
Terminus-2 – GLM-4.7-Flash (Reasoning)43*$0.025No
Terminus-2 – gpt-oss-120b (high)53*$0.026No
Terminus-2 – Kimi K2.7 Code (Reasoning)73*$0.126Yes
Terminus-2 – Qwen 3.8 27B (xhigh)62*$0.066No

Axis Spatial Geospatial Agent Index vs. time per task

Higher and further left is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
  • Most attractive quadrant
  • Pareto line
Axis Spatial Geospatial Agent Index vs. time per taskTerminus-2 – DeepSeek V4 Flash (high): Geospatial Agent Index 69.95, time per task 10.0 min; Terminus-2 – DeepSeek V4 Pro (high): Geospatial Agent Index 72.7, time per task 8.6 min; Terminus-2 – Gemma 4 26B A4B (Reasoning): Geospatial Agent Index 59.63, time per task 8.6 min; Terminus-2 – GLM-4.7-Flash (Reasoning): Geospatial Agent Index 43.05, time per task 9.5 min; Terminus-2 – gpt-oss-120b (high): Geospatial Agent Index 52.72, time per task 4.6 min; Terminus-2 – Kimi K2.7 Code (Reasoning): Geospatial Agent Index 72.62, time per task 6.7 min; Terminus-2 – Qwen 3.8 27B (xhigh): Geospatial Agent Index 62.5, time per task 19.8 minMost attractive quadrant02550751000.0 min6.2 min12.5 min18.8 min25.0 minTime per taskGeospatial Agent IndexDeepSeek V4 Pro (high)Kimi K2.7 Code (Reasoning)DeepSeek V4 Flash (high)Qwen 3.8 27B (xhigh)Gemma 4 26B A4B (Reasoning)gpt-oss-120b (high)GLM-4.7-Flash (Reasoning)
How to read this chart

What execution time is measuring. Agent wall time per attempt, from the agent's first model call to its last action: model thinking, waiting for the provider and running commands. Sandbox start-up and grading are not included.

How to read this chart. Lower is better. An attempt that reaches the task's time limit scores 0 and counts at the limit.

Data table: Axis Spatial Geospatial Agent Index vs. time per task
Axis Spatial Geospatial Agent Index vs. time per task
ModelGeospatial Agent IndexTime per taskOn the Pareto line
Terminus-2 – DeepSeek V4 Flash (high)70*10.0 minNo
Terminus-2 – DeepSeek V4 Pro (high)73*8.6 minYes
Terminus-2 – Gemma 4 26B A4B (Reasoning)60*8.6 minNo
Terminus-2 – GLM-4.7-Flash (Reasoning)43*9.5 minNo
Terminus-2 – gpt-oss-120b (high)53*4.6 minYes
Terminus-2 – Kimi K2.7 Code (Reasoning)73*6.7 minYes
Terminus-2 – Qwen 3.8 27B (xhigh)62*19.8 minNo

Methodology

Each attempt scores 1 or 0. A benchmark score is the average over its tasks of the share of attempts that passed (pass@1). The index is the plain average of the benchmark scores. We run every model ourselves with the same agent, data and limits; the scores are our results on these tasks, not reproductions of the source papers' numbers. Full methodology.

Explore evaluations