Axis Spatial

Climate and time series: geospatial agent capability

Answers that come from a series over many dates or from gridded climate data: daily, seasonal or annual series and statistics across them, such as counts of hot days, change between consecutive dates and fitted trends, and NetCDF climate and Earth-system files.

24 scored tasks are of this kind. All kinds of work and task formats.

Score

Climate and time series score

Average pass@1 over 24 tasks, 0 to 100 · Higher is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
Climate and time series scoreTerminus-2 – Kimi K2.7 Code (Reasoning): 79 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 67 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 62 (partial coverage); Terminus-2 – gpt-oss-120b (high): 62 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 54 (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 48 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 43 (partial coverage)02040608079Terminus-2Kimi K2.7Code(Reasoning)67Terminus-2DeepSeek V4Flash (high)62Terminus-2DeepSeek V4Pro (high)62Terminus-2gpt-oss-120b(high)54Terminus-2GLM-4.7-Flash(Reasoning)48Terminus-2Gemma 4 26BA4B(Reasoning)43Terminus-2Qwen 3.8 27B(xhigh)
How to read this chart

What this metric means. For each of the 24 tasks whose main kind of work is climate and time series, the share of a model's attempts that passed every check; then the plain average over those tasks. Tasks with no decided attempt yet are left out. These scores do not enter the Geospatial Agent Index.

Data table: Climate and time series score
Climate and time series score
ModelCreatorClimate and time series scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek67No pending judgements; coverage incomplete24 of 72 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek62No pending judgements; coverage incomplete24 of 72 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google48No pending judgements; coverage incomplete27 of 72 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai54No pending judgements; coverage incomplete24 of 72 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI62No pending judgements; coverage incomplete24 of 72 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI79No pending judgements; coverage incomplete24 of 72 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)43No pending judgements; coverage incomplete21 of 72 planned attempts

Where the tasks come from

2 other tasks also need this kind of work but count towards another kind, so they are not in this score. How each task was assigned: kinds of work.

Explore kinds of work