Axis Spatial

Z.ai models: geospatial agent results

Geospatial Agent Benchmarks tests 1 model configuration from Z.ai.

Highlights

Results

Axis Spatial Geospatial Agent Index

Equal-weight average of the evaluation scores, 0 to 100 · Higher is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
Axis Spatial Geospatial Agent IndexTerminus-2 – DeepSeek V4 Pro (high): 73* (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 73* (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 70* (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 60* (partial coverage); Terminus-2 – gpt-oss-120b (high): 53* (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 43* (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 62* (partial coverage)025507510073*Terminus-2DeepSeek V4Pro (high)73*Terminus-2Kimi K2.7Code(Reasoning)70*Terminus-2DeepSeek V4Flash (high)60*Terminus-2Gemma 4 26BA4B(Reasoning)53*Terminus-2gpt-oss-120b(high)43*Terminus-2GLM-4.7-Flash(Reasoning)62*Terminus-2Qwen 3.8 27B(xhigh)
Data table: Axis Spatial Geospatial Agent Index
Axis Spatial Geospatial Agent Index
ModelCreatorGeospatial Agent IndexRange while attempts are pendingBenchmarks coveredCoverage (attempts)
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek73*No pending judgements; coverage incomplete4 of 4195 of 870 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI73*No pending judgements; coverage incomplete4 of 4249 of 870 planned attempts
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek70*70 to 704 of 4169 of 870 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google60*60 to 604 of 4145 of 870 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI53*No pending judgements; coverage incomplete4 of 4249 of 870 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai43*43 to 444 of 4143 of 870 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)62*No pending judgements; coverage incomplete3 of 4101 of 870 planned attempts
Z.ai models
ModelGeospatial Agent IndexGeoAgentBenchGeoBenchXEarth-BenchGeoAnalystBenchCost per taskTime per taskContext window
Terminus-2 – GLM-4.7-Flash (Reasoning)43*36464842$0.0259.5 min131,072