Axis Spatial

Capabilities: geospatial agent scores by task format and kind of work

The same attempts as the Geospatial Agent Index, grouped two ways instead of by source benchmark: by task format, what the agent must deliver, and by the kind of geospatial work each task needs. Each of the 290 scored tasks has one format and one main kind of work, so a model's strengths and gaps show across benchmarks. These scores do not enter the index.

By task format

  • Multi-step analysis (151 tasks): carry out a workflow, deliver data files and maps.
  • Single-answer questions (60 tasks): one number, choice or list.
  • Recognising infeasible requests (79 tasks): decline when the data cannot answer.
Score by task format (0 to 100) and the Geospatial Agent Index
ModelMulti-step analysis (151 tasks)Single-answer questions (60 tasks)Recognising infeasible requests (79 tasks)Geospatial Agent Index
Terminus-2 – DeepSeek V4 Pro (high)84454173*
Terminus-2 – Kimi K2.7 Code (Reasoning)79604173*
Terminus-2 – DeepSeek V4 Flash (high)88403570*
Terminus-2 – Gemma 4 26B A4B (Reasoning)78324460*
Terminus-2 – gpt-oss-120b (high)52464753*
Terminus-2 – GLM-4.7-Flash (Reasoning)41473643*
Terminus-2 – Qwen 3.8 27B (xhigh)6427No data62*

Multi-step analysis score

Average pass@1 over 151 tasks, 0 to 100 · Higher is better

  • Partial coverage
Multi-step analysis scoreTerminus-2 – DeepSeek V4 Flash (high): 88 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 84 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 79 (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 78 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 64 (partial coverage); Terminus-2 – gpt-oss-120b (high): 52 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 41 (partial coverage)025507510088DeepSeek V4 Flash (high)84DeepSeek V4 Pro (high)79Kimi K2.7 Code (Reasoning)78Gemma 4 26B A4B (Reasoning)64Qwen 3.8 27B (xhigh)52gpt-oss-120b (high)41GLM-4.7-Flash (Reasoning)
Data table: Multi-step analysis score
Multi-step analysis score
ModelCreatorMulti-step analysis scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek8887 to 8890 of 453 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek84No pending judgements; coverage incomplete101 of 453 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google78No pending judgements; coverage incomplete78 of 453 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai4140 to 4181 of 453 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI52No pending judgements; coverage incomplete128 of 453 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI79No pending judgements; coverage incomplete126 of 453 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)64No pending judgements; coverage incomplete53 of 453 planned attempts

Single-answer questions score

Average pass@1 over 60 tasks, 0 to 100 · Higher is better

  • Partial coverage
Single-answer questions scoreTerminus-2 – Kimi K2.7 Code (Reasoning): 60 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 47 (partial coverage); Terminus-2 – gpt-oss-120b (high): 46 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 45 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 40 (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 32 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 27 (partial coverage)01530456060Kimi K2.7 Code (Reasoning)47GLM-4.7-Flash (Reasoning)46gpt-oss-120b (high)45DeepSeek V4 Pro (high)40DeepSeek V4 Flash (high)32Gemma 4 26B A4B (Reasoning)27Qwen 3.8 27B (xhigh)
Data table: Single-answer questions score
Single-answer questions score
ModelCreatorSingle-answer questions scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek40No pending judgements; coverage incomplete53 of 180 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek45No pending judgements; coverage incomplete55 of 180 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google3232 to 3358 of 180 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai47No pending judgements; coverage incomplete51 of 180 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI46No pending judgements; coverage incomplete59 of 180 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI60No pending judgements; coverage incomplete60 of 180 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)27No pending judgements; coverage incomplete48 of 180 planned attempts

Recognising infeasible requests score

Average pass@1 over 79 tasks, 0 to 100 · Higher is better

  • Partial coverage
Recognising infeasible requests scoreTerminus-2 – gpt-oss-120b (high): 47 (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 44 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 41 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 41 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 36 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 35 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): No data01225385047gpt-oss-120b (high)44Gemma 4 26B A4B (Reasoning)41Kimi K2.7 Code (Reasoning)41DeepSeek V4 Pro (high)36GLM-4.7-Flash (Reasoning)35DeepSeek V4 Flash (high)No dataQwen 3.8 27B (xhigh)
Data table: Recognising infeasible requests score
Recognising infeasible requests score
ModelCreatorRecognising infeasible requests scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek3533 to 3726 of 237 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek41No pending judgements; coverage incomplete39 of 237 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google44No pending judgements; coverage incomplete9 of 237 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai36No pending judgements; coverage incomplete11 of 237 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI47No pending judgements; coverage incomplete62 of 237 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI41No pending judgements; coverage incomplete63 of 237 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)No dataNo data0 of 237 planned attempts

By kind of work

Score by kind of work (0 to 100) and the Geospatial Agent Index
ModelRaster and remote sensing (75 tasks)Vector and overlay (18 tasks)Networks and routing (5 tasks)Climate and time series (24 tasks)Spatial statistics and interpolation (24 tasks)Mapping and cartography (65 tasks)Recognising infeasible requests (79 tasks)Geospatial Agent Index
Terminus-2 – DeepSeek V4 Pro (high)75711006256694173*
Terminus-2 – Kimi K2.7 Code (Reasoning)7478807963704173*
Terminus-2 – DeepSeek V4 Flash (high)66731006773763570*
Terminus-2 – Gemma 4 26B A4B (Reasoning)51791004875834460*
Terminus-2 – gpt-oss-120b (high)52561006230434753*
Terminus-2 – GLM-4.7-Flash (Reasoning)3936205446583643*
Terminus-2 – Qwen 3.8 27B (xhigh)3678674362100No data62*

Raster and remote sensing score

Average pass@1 over 75 tasks, 0 to 100 · Higher is better

  • Partial coverage
Raster and remote sensing scoreTerminus-2 – DeepSeek V4 Pro (high): 75 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 74 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 66 (partial coverage); Terminus-2 – gpt-oss-120b (high): 52 (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 51 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 39 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 36 (partial coverage)02040608075DeepSeek V4 Pro (high)74Kimi K2.7 Code (Reasoning)66DeepSeek V4 Flash (high)52gpt-oss-120b (high)51Gemma 4 26B A4B (Reasoning)39GLM-4.7-Flash (Reasoning)36Qwen 3.8 27B (xhigh)
Data table: Raster and remote sensing score
Raster and remote sensing score
ModelCreatorRaster and remote sensing scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek66No pending judgements; coverage incomplete67 of 225 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek75No pending judgements; coverage incomplete68 of 225 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google5151 to 5167 of 225 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai39No pending judgements; coverage incomplete64 of 225 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI52No pending judgements; coverage incomplete73 of 225 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI74No pending judgements; coverage incomplete74 of 225 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)36No pending judgements; coverage incomplete56 of 225 planned attempts

Vector and overlay score

Average pass@1 over 18 tasks, 0 to 100 · Higher is better

  • Partial coverage
Vector and overlay scoreTerminus-2 – Gemma 4 26B A4B (Reasoning): 79 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 78 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 78 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 73 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 71 (partial coverage); Terminus-2 – gpt-oss-120b (high): 56 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 36 (partial coverage)02040608079Gemma 4 26B A4B (Reasoning)78Kimi K2.7 Code (Reasoning)78Qwen 3.8 27B (xhigh)73DeepSeek V4 Flash (high)71DeepSeek V4 Pro (high)56gpt-oss-120b (high)36GLM-4.7-Flash (Reasoning)
Data table: Vector and overlay score
Vector and overlay score
ModelCreatorVector and overlay scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek73No pending judgements; coverage incomplete15 of 54 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek71No pending judgements; coverage incomplete17 of 54 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google79No pending judgements; coverage incomplete14 of 54 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai36No pending judgements; coverage incomplete14 of 54 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI56No pending judgements; coverage incomplete18 of 54 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI78No pending judgements; coverage incomplete18 of 54 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)78No pending judgements; coverage incomplete9 of 54 planned attempts

Networks and routing score

Average pass@1 over 5 tasks, 0 to 100 · Higher is better

  • Partial coverage
Networks and routing scoreTerminus-2 – DeepSeek V4 Flash (high): 100 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 100 (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 100 (partial coverage); Terminus-2 – gpt-oss-120b (high): 100 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 80 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 67 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 20 (partial coverage)0255075100100DeepSeek V4 Flash (high)100DeepSeek V4 Pro (high)100Gemma 4 26B A4B (Reasoning)100gpt-oss-120b (high)80Kimi K2.7 Code (Reasoning)67Qwen 3.8 27B (xhigh)20GLM-4.7-Flash (Reasoning)
Data table: Networks and routing score
Networks and routing score
ModelCreatorNetworks and routing scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek100No pending judgements; coverage incomplete5 of 15 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek100No pending judgements; coverage incomplete5 of 15 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google100No pending judgements; coverage incomplete4 of 15 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai20No pending judgements; coverage incomplete5 of 15 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI100No pending judgements; coverage incomplete5 of 15 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI80No pending judgements; coverage incomplete5 of 15 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)67No pending judgements; coverage incomplete3 of 15 planned attempts

Climate and time series score

Average pass@1 over 24 tasks, 0 to 100 · Higher is better

  • Partial coverage
Climate and time series scoreTerminus-2 – Kimi K2.7 Code (Reasoning): 79 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 67 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 62 (partial coverage); Terminus-2 – gpt-oss-120b (high): 62 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 54 (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 48 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 43 (partial coverage)02040608079Kimi K2.7 Code (Reasoning)67DeepSeek V4 Flash (high)62DeepSeek V4 Pro (high)62gpt-oss-120b (high)54GLM-4.7-Flash (Reasoning)48Gemma 4 26B A4B (Reasoning)43Qwen 3.8 27B (xhigh)
Data table: Climate and time series score
Climate and time series score
ModelCreatorClimate and time series scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek67No pending judgements; coverage incomplete24 of 72 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek62No pending judgements; coverage incomplete24 of 72 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google48No pending judgements; coverage incomplete27 of 72 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai54No pending judgements; coverage incomplete24 of 72 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI62No pending judgements; coverage incomplete24 of 72 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI79No pending judgements; coverage incomplete24 of 72 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)43No pending judgements; coverage incomplete21 of 72 planned attempts

Spatial statistics and interpolation score

Average pass@1 over 24 tasks, 0 to 100 · Higher is better

  • Partial coverage
Spatial statistics and interpolation scoreTerminus-2 – Gemma 4 26B A4B (Reasoning): 75 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 73 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 63 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): 62 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 56 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 46 (partial coverage); Terminus-2 – gpt-oss-120b (high): 30 (partial coverage)02040608075Gemma 4 26B A4B (Reasoning)73DeepSeek V4 Flash (high)63Kimi K2.7 Code (Reasoning)62Qwen 3.8 27B (xhigh)56DeepSeek V4 Pro (high)46GLM-4.7-Flash (Reasoning)30gpt-oss-120b (high)
Data table: Spatial statistics and interpolation score
Spatial statistics and interpolation score
ModelCreatorSpatial statistics and interpolation scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek73No pending judgements; coverage incomplete15 of 72 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek56No pending judgements; coverage incomplete16 of 72 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google75No pending judgements; coverage incomplete12 of 72 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai46No pending judgements; coverage incomplete13 of 72 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI30No pending judgements; coverage incomplete20 of 72 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI63No pending judgements; coverage incomplete19 of 72 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)62No pending judgements; coverage incomplete8 of 72 planned attempts

Mapping and cartography score

Average pass@1 over 65 tasks, 0 to 100 · Higher is better

  • Partial coverage
Mapping and cartography scoreTerminus-2 – Qwen 3.8 27B (xhigh): 100 (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 83 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 76 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 70 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 69 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 58 (partial coverage); Terminus-2 – gpt-oss-120b (high): 43 (partial coverage)0255075100100Qwen 3.8 27B (xhigh)83Gemma 4 26B A4B (Reasoning)76DeepSeek V4 Flash (high)70Kimi K2.7 Code (Reasoning)69DeepSeek V4 Pro (high)58GLM-4.7-Flash (Reasoning)43gpt-oss-120b (high)
Data table: Mapping and cartography score
Mapping and cartography score
ModelCreatorMapping and cartography scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek7672 to 7817 of 195 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek69No pending judgements; coverage incomplete26 of 195 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google83No pending judgements; coverage incomplete12 of 195 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai5854 to 6212 of 195 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI43No pending judgements; coverage incomplete47 of 195 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI70No pending judgements; coverage incomplete46 of 195 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)100No pending judgements; coverage incomplete4 of 195 planned attempts

Recognising infeasible requests score

Average pass@1 over 79 tasks, 0 to 100 · Higher is better

  • Partial coverage
Recognising infeasible requests scoreTerminus-2 – gpt-oss-120b (high): 47 (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 44 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 41 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 41 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 36 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 35 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): No data01225385047gpt-oss-120b (high)44Gemma 4 26B A4B (Reasoning)41Kimi K2.7 Code (Reasoning)41DeepSeek V4 Pro (high)36GLM-4.7-Flash (Reasoning)35DeepSeek V4 Flash (high)No dataQwen 3.8 27B (xhigh)
Data table: Recognising infeasible requests score
Recognising infeasible requests score
ModelCreatorRecognising infeasible requests scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek3533 to 3726 of 237 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek41No pending judgements; coverage incomplete39 of 237 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google44No pending judgements; coverage incomplete9 of 237 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai36No pending judgements; coverage incomplete11 of 237 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI47No pending judgements; coverage incomplete62 of 237 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI41No pending judgements; coverage incomplete63 of 237 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)No dataNo data0 of 237 planned attempts

All kinds of work

How tasks are assigned

The format follows from what each task must deliver; the kind of work is the one its checked result depends on most. Rules and borderline tasks.