Axis Spatial

Image track: geospatial questions answered from images

Earth observation is largely imagery. This separate track asks 272 questions whose answer needs the model to look at satellite or aerial images. It is not part of the Geospatial Agent Index, and its two protocols are shown as separate tables, not one ranking.

Ongoing evaluation: results update as runs complete. Last updated 10 October 2026.

Look and answer

One model call per question: the images and the question go in, the answer comes out. No tools and no code.

Look and answer: score by source (0 to 100)
ModelEarth-Bench-V2 images (100 questions)ThinkGeo (172 questions)Image track scoreQuestions decidedNot supported by the routeCost per questionTime per question
Kimi K2.7 Code (always on)613246123 of 1360$0.0605.0 min
Kimi K2.6 (always on)592944131 of 1360$0.0699.5 min
GLM-5.3-Flash (max) (14 questions not supported)533544122 of 13614$0.00734.0 min
Gemma 4 26B A4B (thinking on)61234288 of 1360$0.00356.0 min
Mistral Small 3.1 24B (Non-reasoning)541233136 of 1360$0.00110.2 min

Look and work

An agent views the image files and may run Python before answering, in an isolated sandbox. Its traces are reviewed for reward hacking as in the main index.

Look and work: score by source (0 to 100)
Agent and modelEarth-Bench-V2 images (100 questions)ThinkGeo (172 questions)Image track scoreQuestions decidedNot supported by the routeCost per questionTime per question
Claude Code – Claude Opus 5.5 (medium)703151136 of 1360$0.0870.3 min
Codex – GPT-6.1 Sol (medium)603346136 of 1360$0.0310.7 min
Claude Code – Claude Sonnet 5.5 (medium)642344136 of 1360$0.0360.2 min
Claude Code – Claude Haiku 5.5 (medium)582441136 of 1360$0.00440.3 min
Codex – GPT-6 Luna (xhigh)441731136 of 1360$0.00501.8 min

How to read these results

  • Each question is scored pass or fail by code against its answer key; the image track score is the plain average of the two source scores. Each question has 1 attempt for now; more will be added and averaged in.
  • "Not supported by the route" counts questions the model's serving route refused because of its input, for example more images than it accepts. They are left out of the score and counted here.
  • Sources: the image questions of Earth-Bench-V2 (dataset LvZTTTT/Earth-Bench-V2 at revision 5115836; Earth-Agent, arXiv 2509.23141; Earth-Agent-Pro, arXiv 2609.12533; annotations CC BY-NC 4.0) and ThinkGeo (MBZUAI/ThinkGeo, CC BY 4.0, imagery under its source datasets' terms). This non-commercial site publishes scores only, never questions or images.
  • Methodology describes both protocols.