Axis Spatial

Leaderboard: AI models as geospatial agents

Every model configuration ranked by the Geospatial Agent Index, with its score on each benchmark, cost and time per task, and its specification. See the overview for charts.

Ongoing evaluation: results update as runs complete. Last updated 9 October 2026.

Highlights

All models

Overall

Ranked: entries with a score on every benchmark. An asterisk marks partial coverage.
RankModelCreatorGeospatial Agent IndexGeoAgentBenchGeoBenchXEarth-BenchGeoAnalystBenchCost per taskTime per taskTokens per attemptAttempts decidedContext windowPrice per 1M input / output tokensImage inputFurther analysis
1Terminus-2 – DeepSeek V4 Pro (high)DeepSeek73*984548100$0.2618.6 min197k1951,048,576$1.32 / $3.96NoModel
2Terminus-2 – Kimi K2.7 Code (Reasoning)Moonshot AI73*90546284$0.1266.7 min309k249262,144$0.95 / $4.00YesModel
3Terminus-2 – DeepSeek V4 Flash (high)DeepSeek70*964440100$0.06010.0 min300k1691,048,576$0.44 / $1.32NoModel
4Terminus-2 – Gemma 4 26B A4B (Reasoning)Google60*86553068$0.0128.6 min89k145256,000$0.10 / $0.30YesModel
5Terminus-2 – gpt-oss-120b (high)OpenAI53*70424653$0.0264.6 min59k249128,000$0.35 / $0.75NoModel
6Terminus-2 – GLM-4.7-Flash (Reasoning)Z.ai43*36464842$0.0259.5 min332k143131,072$0.06 / $0.40NoModel
Not yet ranked: missing benchmark scores. Each entry's index covers only the benchmarks it has scores for.
RankModelCreatorGeospatial Agent IndexGeoAgentBenchGeoBenchXEarth-BenchGeoAnalystBenchCost per taskTime per taskTokens per attemptAttempts decidedContext windowPrice per 1M input / output tokensImage inputFurther analysis
–Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)62*60No data27100$0.06619.8 min90k101262,144$0.45 / $3.20YesModel

By task format

Score by task format (0 to 100) and the Geospatial Agent Index
ModelMulti-step analysis (151 tasks)Single-answer questions (60 tasks)Recognising infeasible requests (79 tasks)Geospatial Agent Index
Terminus-2 – DeepSeek V4 Pro (high)84454173*
Terminus-2 – Kimi K2.7 Code (Reasoning)79604173*
Terminus-2 – DeepSeek V4 Flash (high)88403570*
Terminus-2 – Gemma 4 26B A4B (Reasoning)78324460*
Terminus-2 – gpt-oss-120b (high)52464753*
Terminus-2 – GLM-4.7-Flash (Reasoning)41473643*
Terminus-2 – Qwen 3.8 27B (xhigh)6427No data62*

By kind of work

Score by kind of work (0 to 100) and the Geospatial Agent Index
ModelRaster and remote sensing (75 tasks)Vector and overlay (18 tasks)Networks and routing (5 tasks)Climate and time series (24 tasks)Spatial statistics and interpolation (24 tasks)Mapping and cartography (65 tasks)Recognising infeasible requests (79 tasks)Geospatial Agent Index
Terminus-2 – DeepSeek V4 Pro (high)75711006256694173*
Terminus-2 – Kimi K2.7 Code (Reasoning)7478807963704173*
Terminus-2 – DeepSeek V4 Flash (high)66731006773763570*
Terminus-2 – Gemma 4 26B A4B (Reasoning)51791004875834460*
Terminus-2 – gpt-oss-120b (high)52561006230434753*
Terminus-2 – GLM-4.7-Flash (Reasoning)3936205446583643*
Terminus-2 – Qwen 3.8 27B (xhigh)3678674362100No data62*