Image track: geospatial questions answered from images
Earth observation is largely imagery. This separate track asks 272 questions whose answer needs the model to look at satellite or aerial images. It is not part of the Geospatial Agent Index, and its two protocols are shown as separate tables, not one ranking.
Ongoing evaluation: results update as runs complete. Last updated 10 October 2026.
Look and answer
One model call per question: the images and the question go in, the answer comes out. No tools and no code.
| Model | Earth-Bench-V2 images (100 questions) | ThinkGeo (172 questions) | Image track score | Questions decided | Not supported by the route | Cost per question | Time per question |
|---|---|---|---|---|---|---|---|
| Kimi K2.7 Code (always on) | 61 | 32 | 46 | 123 of 136 | 0 | $0.060 | 5.0 min |
| Kimi K2.6 (always on) | 59 | 29 | 44 | 131 of 136 | 0 | $0.069 | 9.5 min |
| GLM-5.3-Flash (max) (14 questions not supported) | 53 | 35 | 44 | 122 of 136 | 14 | $0.0073 | 4.0 min |
| Gemma 4 26B A4B (thinking on) | 61 | 23 | 42 | 88 of 136 | 0 | $0.0035 | 6.0 min |
| Mistral Small 3.1 24B (Non-reasoning) | 54 | 12 | 33 | 136 of 136 | 0 | $0.0011 | 0.2 min |
Look and work
An agent views the image files and may run Python before answering, in an isolated sandbox. Its traces are reviewed for reward hacking as in the main index.
| Agent and model | Earth-Bench-V2 images (100 questions) | ThinkGeo (172 questions) | Image track score | Questions decided | Not supported by the route | Cost per question | Time per question |
|---|---|---|---|---|---|---|---|
| Claude Code – Claude Opus 5.5 (medium) | 70 | 31 | 51 | 136 of 136 | 0 | $0.087 | 0.3 min |
| Codex – GPT-6.1 Sol (medium) | 60 | 33 | 46 | 136 of 136 | 0 | $0.031 | 0.7 min |
| Claude Code – Claude Sonnet 5.5 (medium) | 64 | 23 | 44 | 136 of 136 | 0 | $0.036 | 0.2 min |
| Claude Code – Claude Haiku 5.5 (medium) | 58 | 24 | 41 | 136 of 136 | 0 | $0.0044 | 0.3 min |
| Codex – GPT-6 Luna (xhigh) | 44 | 17 | 31 | 136 of 136 | 0 | $0.0050 | 1.8 min |
How to read these results
- Each question is scored pass or fail by code against its answer key; the image track score is the plain average of the two source scores. Each question has 1 attempt for now; more will be added and averaged in.
- "Not supported by the route" counts questions the model's serving route refused because of its input, for example more images than it accepts. They are left out of the score and counted here.
- Sources: the image questions of Earth-Bench-V2 (dataset LvZTTTT/Earth-Bench-V2 at revision 5115836; Earth-Agent, arXiv 2509.23141; Earth-Agent-Pro, arXiv 2609.12533; annotations CC BY-NC 4.0) and ThinkGeo (MBZUAI/ThinkGeo, CC BY 4.0, imagery under its source datasets' terms). This non-commercial site publishes scores only, never questions or images.
- Methodology describes both protocols.
