Axis Spatial

Terminus-2 – DeepSeek V4 Pro (high): geospatial agent results

Terminus-2 – DeepSeek V4 Pro (high) by DeepSeek, run under the Terminus-2 agent: its Geospatial Agent Index, score on each benchmark and each kind of geospatial work, cost, time and token use.

Summary

Geospatial Agent Index
73* (rank 1 of 6)
Cost per task
$0.261
Time per task
8.6 min
Tokens per attempt
197k
Turns per attempt
12.6
Attempts decided
195

Score by benchmark

Terminus-2 – DeepSeek V4 Pro (high): score by benchmark

Average pass@1, 0 to 100 · Higher is better

Terminus-2 – DeepSeek V4 Pro (high): score by benchmarkGeoAgentBench: 98 (partial coverage); GeoBenchX: 45 (partial coverage); Earth-Bench: 48 (partial coverage); GeoAnalystBench: 100 (partial coverage)025507510098GeoAgentBench45GeoBenchX48Earth-Bench100GeoAnalystBench
Data table: Terminus-2 – DeepSeek V4 Pro (high): score by benchmark
Terminus-2 – DeepSeek V4 Pro (high): by benchmark
BenchmarkScoreTasks with a decided attemptCost per taskTime per task
GeoAgentBench9850 of 50$0.0915.5 min
GeoBenchX4578 of 173$0.2847.3 min
Earth-Bench4848 of 48$0.46815.9 min
GeoAnalystBench10019 of 19$0.0873.3 min

Score by task format and kind of work

Terminus-2 – DeepSeek V4 Pro (high): score by task format

Average pass@1 over the tasks of each task format, 0 to 100 · Higher is better

Terminus-2 – DeepSeek V4 Pro (high): score by task formatMulti-step analysis: 84 (partial coverage); Single-answer questions: 45 (partial coverage); Recognising infeasible requests: 41 (partial coverage)025507510084Multi-stepanalysis45Single-answerquestions41Recognisinginfeasiblerequests
How to read this chart

What this shows. The same attempts as the index, grouped by task format. Groups with few tasks are less certain. About these breakdowns.

Data table: Terminus-2 – DeepSeek V4 Pro (high): score by task format
Terminus-2 – DeepSeek V4 Pro (high): by task format
Task formatScoreTasks with a decided attempt
Multi-step analysis84101 of 151
Single-answer questions4555 of 60
Recognising infeasible requests4139 of 79

Terminus-2 – DeepSeek V4 Pro (high): score by kind of work

Average pass@1 over the tasks of each kind of work, 0 to 100 · Higher is better

Terminus-2 – DeepSeek V4 Pro (high): score by kind of workRaster and remote sensing: 75 (partial coverage); Vector and overlay: 71 (partial coverage); Networks and routing: 100 (partial coverage); Climate and time series: 62 (partial coverage); Spatial statistics and interpolation: 56 (partial coverage); Mapping and cartography: 69 (partial coverage); Recognising infeasible requests: 41 (partial coverage)025507510075Raster andremotesensing71Vector andoverlay100Networks androuting62Climate andtime series56Spatialstatisticsand69Mapping andcartography41Recognisinginfeasiblerequests
How to read this chart

What this shows. The same attempts as the index, grouped by kind of work. Groups with few tasks are less certain. About these breakdowns.

Data table: Terminus-2 – DeepSeek V4 Pro (high): score by kind of work
Terminus-2 – DeepSeek V4 Pro (high): by kind of work
Kind of workScoreTasks with a decided attempt
Raster and remote sensing7568 of 75
Vector and overlay7117 of 18
Networks and routing1005 of 5
Climate and time series6224 of 24
Spatial statistics and interpolation5616 of 24
Mapping and cartography6926 of 65
Recognising infeasible requests4139 of 79

Compared with other models

Axis Spatial Geospatial Agent Index

Equal-weight average of the evaluation scores, 0 to 100 · Higher is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
Axis Spatial Geospatial Agent IndexTerminus-2 – DeepSeek V4 Pro (high): 73* (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 73* (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 70* (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 60* (partial coverage); Terminus-2 – gpt-oss-120b (high): 53* (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 43* (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 62* (partial coverage)025507510073*Terminus-2DeepSeek V4Pro (high)73*Terminus-2Kimi K2.7Code(Reasoning)70*Terminus-2DeepSeek V4Flash (high)60*Terminus-2Gemma 4 26BA4B(Reasoning)53*Terminus-2gpt-oss-120b(high)43*Terminus-2GLM-4.7-Flash(Reasoning)62*Terminus-2Qwen 3.8 27B(xhigh)
How to read this chart

What this metric means. The equal-weight average of the 4 evaluation scores. Each evaluation score is the average pass@1 over that benchmark's tasks, from 0 to 100.

Not yet ranked. Entries without a score on every benchmark come after the ranked entries and are not ranked: their index averages only the benchmarks they have.

Partial coverage. A hatched bar, and an asterisk in tables, marks an entry that does not yet have results for every planned attempt. While attempts await review, the table gives the range the score could still reach.

Data table: Axis Spatial Geospatial Agent Index
Axis Spatial Geospatial Agent Index
ModelCreatorGeospatial Agent IndexRange while attempts are pendingBenchmarks coveredCoverage (attempts)
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek73*No pending judgements; coverage incomplete4 of 4195 of 870 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI73*No pending judgements; coverage incomplete4 of 4249 of 870 planned attempts
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek70*70 to 704 of 4169 of 870 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google60*60 to 604 of 4145 of 870 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI53*No pending judgements; coverage incomplete4 of 4249 of 870 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai43*43 to 444 of 4143 of 870 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)62*No pending judgements; coverage incomplete3 of 4101 of 870 planned attempts

Axis Spatial Geospatial Agent Index vs. cost per task

Higher and further left is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
  • Most attractive quadrant
  • Pareto line
Axis Spatial Geospatial Agent Index vs. cost per taskTerminus-2 – DeepSeek V4 Flash (high): Geospatial Agent Index 69.95, cost per task (usd) $0.060; Terminus-2 – DeepSeek V4 Pro (high): Geospatial Agent Index 72.7, cost per task (usd) $0.261; Terminus-2 – Gemma 4 26B A4B (Reasoning): Geospatial Agent Index 59.63, cost per task (usd) $0.012; Terminus-2 – GLM-4.7-Flash (Reasoning): Geospatial Agent Index 43.05, cost per task (usd) $0.025; Terminus-2 – gpt-oss-120b (high): Geospatial Agent Index 52.72, cost per task (usd) $0.026; Terminus-2 – Kimi K2.7 Code (Reasoning): Geospatial Agent Index 72.62, cost per task (usd) $0.126; Terminus-2 – Qwen 3.8 27B (xhigh): Geospatial Agent Index 62.5, cost per task (usd) $0.066Most attractive quadrant0255075100$0.010$0.020$0.050$0.100$0.200$0.500$1.00Cost per task (USD) (log scale)Geospatial Agent IndexDeepSeek V4 Pro (high)Kimi K2.7 Code (Reasoning)DeepSeek V4 Flash (high)Qwen 3.8 27B (xhigh)Gemma 4 26B A4B (Reasoning)gpt-oss-120b (high)GLM-4.7-Flash (Reasoning)
Data table: Axis Spatial Geospatial Agent Index vs. cost per task
Axis Spatial Geospatial Agent Index vs. cost per task
ModelGeospatial Agent IndexCost per task (USD)On the Pareto line
Terminus-2 – DeepSeek V4 Flash (high)70*$0.060Yes
Terminus-2 – DeepSeek V4 Pro (high)73*$0.261Yes
Terminus-2 – Gemma 4 26B A4B (Reasoning)60*$0.012Yes
Terminus-2 – GLM-4.7-Flash (Reasoning)43*$0.025No
Terminus-2 – gpt-oss-120b (high)53*$0.026No
Terminus-2 – Kimi K2.7 Code (Reasoning)73*$0.126Yes
Terminus-2 – Qwen 3.8 27B (xhigh)62*$0.066No

Head-to-head comparisons

Specification and settings

Developer
DeepSeek
Context window
1,048,576 tokens
Image input
No
Reasoning setting
high
Temperature
0.6
Maximum output tokens
65,536
Input price per 1M tokens
$1.32
Cached input price per 1M tokens
$0.044
Output price per 1M tokens
$3.96

Why this model was chosen: model selection.