Axis Spatial

Terminus-2 – DeepSeek V4 Pro (high) vs. Terminus-2 – GLM-4.7-Flash (thinking on): geospatial agent comparison

Comparison between Terminus-2 – DeepSeek V4 Pro (high) and Terminus-2 – GLM-4.7-Flash (thinking on) on the Geospatial Agent Index, each benchmark, cost, time and token use, from the same tasks, agent and limits. Methodology.

Side by side

Terminus-2 – DeepSeek V4 Pro (high) vs. Terminus-2 – GLM-4.7-Flash (thinking on)
MetricTerminus-2 – DeepSeek V4 Pro (high)Terminus-2 – GLM-4.7-Flash (thinking on)
Geospatial Agent Index73*41*
GeoAgentBench score9836
GeoBenchX score4639
Earth-Bench score4845
GeoAnalystBench score10042
Raster and remote sensing score7436
Vector and overlay score7136
Networks and routing score10020
Climate and time series score6555
Spatial statistics and interpolation score5340
Mapping and cartography score6954
Recognising infeasible requests score4236
Cost per task$0.265$0.027
Time per task8.6 min10.0 min
Tokens per attempt201k357k
Turns per attempt12.816.6
Attempts decided209176
Input price per 1M tokens$1.32$0.06
Output price per 1M tokens$3.96$0.40
Context window1,048,576131,072
Image inputNoNo
DeveloperDeepSeekZ.ai

Charts

Axis Spatial Geospatial Agent Index

Equal-weight average of the evaluation scores, 0 to 100 · Higher is better

  • Partial coverage
Axis Spatial Geospatial Agent IndexTerminus-2 – DeepSeek V4 Pro (high): 73* (partial coverage); Terminus-2 – GLM-4.7-Flash (thinking on): 41* (partial coverage)025507510073*DeepSeek V4 Pro (high)41*GLM-4.7-Flash (thinking on)
Data table: Axis Spatial Geospatial Agent Index
Axis Spatial Geospatial Agent Index
ModelCreatorGeospatial Agent IndexRange while attempts are pendingBenchmarks coveredCoverage (attempts)
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek73*No pending judgements; coverage incomplete4 of 4209 of 870 planned attempts
Terminus-2 – GLM-4.7-Flash (thinking on)Z.ai41*No pending judgements; coverage incomplete4 of 4176 of 870 planned attempts

Cost per task

Average cost per task (USD) · Lower is better

  • Partial coverage
Cost per taskTerminus-2 – GLM-4.7-Flash (thinking on): $0.027 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): $0.265 (partial coverage)$0$0.075$0.150$0.225$0.300$0.027GLM-4.7-Flash (thinking on)$0.265DeepSeek V4 Pro (high)
Data table: Cost per task
Cost per task
ModelCreatorCost per taskCoverage (attempts)
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek$0.265209 of 870 planned attempts
Terminus-2 – GLM-4.7-Flash (thinking on)Z.ai$0.027176 of 870 planned attempts

Time per task

Average agent wall time per task · Lower is better

  • Partial coverage
Time per taskTerminus-2 – DeepSeek V4 Pro (high): 8.6 min (partial coverage); Terminus-2 – GLM-4.7-Flash (thinking on): 10.0 min (partial coverage)0m3m7m10m13m8.6 minDeepSeek V4 Pro (high)10.0 minGLM-4.7-Flash (thinking on)
Data table: Time per task
Time per task
ModelCreatorTime per taskCoverage (attempts)
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek8.6 min209 of 870 planned attempts
Terminus-2 – GLM-4.7-Flash (thinking on)Z.ai10.0 min176 of 870 planned attempts

Frequently asked questions

Which scores higher on the Geospatial Agent Index, Terminus-2 – DeepSeek V4 Pro (high) or Terminus-2 – GLM-4.7-Flash (thinking on)?
Terminus-2 – DeepSeek V4 Pro (high), at 73 against 41 for Terminus-2 – GLM-4.7-Flash (thinking on). Terminus-2 – DeepSeek V4 Pro (high): 209 of 870 planned attempts. Terminus-2 – GLM-4.7-Flash (thinking on): 176 of 870 planned attempts. Coverage is partial; scores can change as more attempts are decided.
Which is cheaper per task, Terminus-2 – DeepSeek V4 Pro (high) or Terminus-2 – GLM-4.7-Flash (thinking on)?
Terminus-2 – GLM-4.7-Flash (thinking on), at $0.027 against $0.265 for Terminus-2 – DeepSeek V4 Pro (high).
Which is faster per task, Terminus-2 – DeepSeek V4 Pro (high) or Terminus-2 – GLM-4.7-Flash (thinking on)?
Terminus-2 – DeepSeek V4 Pro (high), at 8.6 min against 10.0 min for Terminus-2 – GLM-4.7-Flash (thinking on).
Which has the larger context window, Terminus-2 – DeepSeek V4 Pro (high) or Terminus-2 – GLM-4.7-Flash (thinking on)?
Terminus-2 – DeepSeek V4 Pro (high), at 1,048,576 tokens against 131,072 tokens for Terminus-2 – GLM-4.7-Flash (thinking on).

More: Terminus-2 – DeepSeek V4 Pro (high) · Terminus-2 – GLM-4.7-Flash (thinking on) · all comparisons