Axis Spatial

Recognising infeasible requests: geospatial agent capability

GeoBenchX tasks that cannot be done with the data provided, where the only passing answer is to reject the task. Every GeoBenchX task asks for the same output files, so the instruction does not reveal which tasks these are.

79 scored tasks are of this kind. All kinds of work and task formats.

Score

Recognising infeasible requests score

Average pass@1 over 79 tasks, 0 to 100 · Higher is better

  • DeepSeek
  • Google
  • Z.ai
  • OpenAI
  • Moonshot AI
  • Alibaba (Qwen)
  • Partial coverage
Recognising infeasible requests scoreTerminus-2 – gpt-oss-120b (high): 47 (partial coverage); Terminus-2 – Gemma 4 26B A4B (Reasoning): 44 (partial coverage); Terminus-2 – Kimi K2.7 Code (Reasoning): 41 (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 41 (partial coverage); Terminus-2 – GLM-4.7-Flash (Reasoning): 36 (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 35 (partial coverage); Terminus-2 – Qwen 3.8 27B (xhigh): No data01225385047Terminus-2gpt-oss-120b(high)44Terminus-2Gemma 4 26BA4B(Reasoning)41Terminus-2Kimi K2.7Code(Reasoning)41Terminus-2DeepSeek V4Pro (high)36Terminus-2GLM-4.7-Flash(Reasoning)35Terminus-2DeepSeek V4Flash (high)No dataTerminus-2Qwen 3.8 27B(xhigh)
How to read this chart

What this metric means. For each of the 79 tasks whose main kind of work is recognising infeasible requests, the share of a model's attempts that passed every check; then the plain average over those tasks. Tasks with no decided attempt yet are left out. These scores do not enter the Geospatial Agent Index.

Data table: Recognising infeasible requests score
Recognising infeasible requests score
ModelCreatorRecognising infeasible requests scoreRange while attempts are pendingCoverage (attempts)
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek3533 to 3726 of 237 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek41No pending judgements; coverage incomplete39 of 237 planned attempts
Terminus-2 – Gemma 4 26B A4B (Reasoning)Google44No pending judgements; coverage incomplete9 of 237 planned attempts
Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai36No pending judgements; coverage incomplete11 of 237 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI47No pending judgements; coverage incomplete62 of 237 planned attempts
Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI41No pending judgements; coverage incomplete63 of 237 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)No dataNo data0 of 237 planned attempts

Where the tasks come from

5 other tasks also need this kind of work but count towards another kind, so they are not in this score. How each task was assigned: kinds of work.

Explore kinds of work