Axis Spatial

Spatial statistics and interpolation: geospatial agent capability

Statistical models of space: kriging and other interpolation, kernel density and heatmaps, local Moran's I hot spots, regression including geographically weighted regression, point-pattern functions and clustering.

24 scored tasks are of this kind. All kinds of work and task formats.

Score

Spatial statistics and interpolation score

Average pass@1 over 24 tasks, 0 to 100 · Higher is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
Spatial statistics and interpolation scoreTerminus-2 – Gemma 4 26B A4B (Reasoning): 75 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 73 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 63 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 62 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 56 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 46 (partial coverage); Terminus-2 – gpt-oss-120b (high): 30 (partial coverage)02040608075Terminus-2Gemma 4 26BA4B(Reasoning)73Terminus-2DeepSeek V4Flash (high)63Terminus-2Kimi K2.7Code(Reasoning)62Terminus-2Qwen 3.8 27B(xhigh)56Terminus-2DeepSeek V4Pro (high)46Terminus-2GLM-4.7-Flash(Reasoning)30Terminus-2gpt-oss-120b(high)
How to read this chart

What this metric means. For each of the 24 tasks whose main kind of work is spatial statistics and interpolation, the share of a model's attempts that passed every check; then the plain average over those tasks. Tasks with no decided attempt yet are left out. These scores do not enter the Geospatial Agent Index.

Data table: Spatial statistics and interpolation score
Spatial statistics and interpolation score
ModelCreatorSpatial statistics and interpolation scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek73No pending judgements; coverage incomplete15 of 72 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek56No pending judgements; coverage incomplete16 of 72 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google75No pending judgements; coverage incomplete12 of 72 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai46No pending judgements; coverage incomplete13 of 72 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI30No pending judgements; coverage incomplete20 of 72 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI63No pending judgements; coverage incomplete19 of 72 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)62No pending judgements; coverage incomplete8 of 72 planned attempts

Where the tasks come from

6 other tasks also need this kind of work but count towards another kind, so they are not in this score. How each task was assigned: kinds of work.

Explore kinds of work