Axis Spatial

Raster and remote sensing: geospatial agent capability

Computing over gridded data: satellite retrievals such as land surface temperature, NDVI and burn ratios, reclassifying and combining rasters, terrain and hydrology from elevation models, and values sampled from a raster. A comparison of two dates belongs here.

75 scored tasks are of this kind. All kinds of work and task formats.

Score

Raster and remote sensing score

Average pass@1 over 75 tasks, 0 to 100 · Higher is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
Raster and remote sensing scoreTerminus-2 – DeepSeek V4 Pro (high): 75 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 74 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 66 (partial coverage); Terminus-2 – gpt-oss-120b (high): 52 (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 51 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 39 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 36 (partial coverage)02040608075Terminus-2DeepSeek V4Pro (high)74Terminus-2Kimi K2.7Code(Reasoning)66Terminus-2DeepSeek V4Flash (high)52Terminus-2gpt-oss-120b(high)51Terminus-2Gemma 4 26BA4B(Reasoning)39Terminus-2GLM-4.7-Flash(Reasoning)36Terminus-2Qwen 3.8 27B(xhigh)
How to read this chart

What this metric means. For each of the 75 tasks whose main kind of work is raster and remote sensing, the share of a model's attempts that passed every check; then the plain average over those tasks. Tasks with no decided attempt yet are left out. These scores do not enter the Geospatial Agent Index.

Data table: Raster and remote sensing score
Raster and remote sensing score
ModelCreatorRaster and remote sensing scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek66No pending judgements; coverage incomplete67 of 225 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek75No pending judgements; coverage incomplete68 of 225 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google5151 to 5167 of 225 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai39No pending judgements; coverage incomplete64 of 225 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI52No pending judgements; coverage incomplete73 of 225 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI74No pending judgements; coverage incomplete74 of 225 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)36No pending judgements; coverage incomplete56 of 225 planned attempts

Where the tasks come from

22 other tasks also need this kind of work but count towards another kind, so they are not in this score. How each task was assigned: kinds of work.

Explore kinds of work