Axis Spatial

Mapping and cartography: geospatial agent capability

Tasks where the map is the main deliverable and the analysis behind it is light: joining tables to boundaries, filtering, then a choropleth, bivariate or point map. The map is graded by an AI judge against a private checklist.

65 scored tasks are of this kind. All kinds of work and task formats.

Score

Mapping and cartography score

Average pass@1 over 65 tasks, 0 to 100 · Higher is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
Mapping and cartography scoreTerminus-2 – Qwen 3.8 27B (xhigh): 100 (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 83 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 76 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 70 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 69 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 58 (partial coverage); Terminus-2 – gpt-oss-120b (high): 43 (partial coverage)0255075100100Terminus-2Qwen 3.8 27B(xhigh)83Terminus-2Gemma 4 26BA4B(Reasoning)76Terminus-2DeepSeek V4Flash (high)70Terminus-2Kimi K2.7Code(Reasoning)69Terminus-2DeepSeek V4Pro (high)58Terminus-2GLM-4.7-Flash(Reasoning)43Terminus-2gpt-oss-120b(high)
How to read this chart

What this metric means. For each of the 65 tasks whose main kind of work is mapping and cartography, the share of a model's attempts that passed every check; then the plain average over those tasks. Tasks with no decided attempt yet are left out. These scores do not enter the Geospatial Agent Index.

Data table: Mapping and cartography score
Mapping and cartography score
ModelCreatorMapping and cartography scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek7672 to 7817 of 195 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek69No pending judgements; coverage incomplete26 of 195 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google83No pending judgements; coverage incomplete12 of 195 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai5854 to 6212 of 195 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI43No pending judgements; coverage incomplete47 of 195 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI70No pending judgements; coverage incomplete46 of 195 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)100No pending judgements; coverage incomplete4 of 195 planned attempts

Where the tasks come from

39 other tasks also need this kind of work but count towards another kind, so they are not in this score. How each task was assigned: kinds of work.

Explore kinds of work