Axis Spatial

Vector and overlay: geospatial agent capability

Geometry operations on points, lines and polygons: buffers, spatial selection and joins, intersection and difference, dissolve, reprojection, and counting points in polygons or grids.

18 scored tasks are of this kind. All kinds of work and task formats.

Score

Vector and overlay score

Average pass@1 over 18 tasks, 0 to 100 · Higher is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
Vector and overlay scoreTerminus-2 – Gemma 4 26B A4B (Reasoning): 79 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 78 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 78 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 73 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 71 (partial coverage); Terminus-2 – gpt-oss-120b (high): 56 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 36 (partial coverage)02040608079Terminus-2Gemma 4 26BA4B(Reasoning)78Terminus-2Kimi K2.7Code(Reasoning)78Terminus-2Qwen 3.8 27B(xhigh)73Terminus-2DeepSeek V4Flash (high)71Terminus-2DeepSeek V4Pro (high)56Terminus-2gpt-oss-120b(high)36Terminus-2GLM-4.7-Flash(Reasoning)
How to read this chart

What this metric means. For each of the 18 tasks whose main kind of work is vector and overlay, the share of a model's attempts that passed every check; then the plain average over those tasks. Tasks with no decided attempt yet are left out. These scores do not enter the Geospatial Agent Index.

Data table: Vector and overlay score
Vector and overlay score
ModelCreatorVector and overlay scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek73No pending judgements; coverage incomplete15 of 54 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek71No pending judgements; coverage incomplete17 of 54 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google79No pending judgements; coverage incomplete14 of 54 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai36No pending judgements; coverage incomplete14 of 54 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI56No pending judgements; coverage incomplete18 of 54 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI78No pending judgements; coverage incomplete18 of 54 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)78No pending judgements; coverage incomplete9 of 54 planned attempts

Where the tasks come from

32 other tasks also need this kind of work but count towards another kind, so they are not in this score. How each task was assigned: kinds of work.

Explore kinds of work