Axis Spatial

Networks and routing: geospatial agent capability

Work that needs a graph: shortest and least-time routes, service areas, and origin-destination costs and flows over a road network.

5 scored tasks are of this kind. All kinds of work and task formats.

Score

Networks and routing score

Average pass@1 over 5 tasks, 0 to 100 · Higher is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
Networks and routing scoreTerminus-2 – DeepSeek V4 Flash (high): 100 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 100 (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 100 (partial coverage); Terminus-2 – gpt-oss-120b (high): 100 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 80 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 67 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 20 (partial coverage)0255075100100Terminus-2DeepSeek V4Flash (high)100Terminus-2DeepSeek V4Pro (high)100Terminus-2Gemma 4 26BA4B(Reasoning)100Terminus-2gpt-oss-120b(high)80Terminus-2Kimi K2.7Code(Reasoning)67Terminus-2Qwen 3.8 27B(xhigh)20Terminus-2GLM-4.7-Flash(Reasoning)
How to read this chart

What this metric means. For each of the 5 tasks whose main kind of work is networks and routing, the share of a model's attempts that passed every check; then the plain average over those tasks. Tasks with no decided attempt yet are left out. These scores do not enter the Geospatial Agent Index.

Few tasks. Only 5 tasks are of this kind, so one task moves the score by a lot. Compare models here with care.

Data table: Networks and routing score
Networks and routing score
ModelCreatorNetworks and routing scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek100No pending judgements; coverage incomplete5 of 15 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek100No pending judgements; coverage incomplete5 of 15 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google100No pending judgements; coverage incomplete4 of 15 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai20No pending judgements; coverage incomplete5 of 15 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI100No pending judgements; coverage incomplete5 of 15 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI80No pending judgements; coverage incomplete5 of 15 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)67No pending judgements; coverage incomplete3 of 15 planned attempts

Where the tasks come from

2 other tasks also need this kind of work but count towards another kind, so they are not in this score. How each task was assigned: kinds of work.

Explore kinds of work