Axis Spatial

Codex – GPT-6 Luna (xhigh): geospatial agent results

Codex – GPT-6 Luna (xhigh) by OpenAI, run under the Terminus-2 agent: its Geospatial Agent Index, score on each benchmark and each kind of geospatial work, cost, time and token use.

Summary

Geospatial Agent Index
77* (partial: 1.9 of 3 attempts) (rank 9 of 30)
Cost per task
$0.0092
Time per task
3.4 min
Tokens per attempt
236k
Turns per attempt
9.8
Attempts decided
512

Score by benchmark

Codex – GPT-6 Luna (xhigh): score by benchmark

Average pass@1, 0 to 100 · Higher is better

Codex – GPT-6 Luna (xhigh): score by benchmarkGeoAgentBench: 99 (partial coverage); GeoBenchX: 60 (partial coverage); Earth-Bench: 55 (partial coverage); GeoAnalystBench: 95 (partial coverage)025507510099GeoAgentBench60GeoBenchX55Earth-Bench95GeoAnalystBench
Data table: Codex – GPT-6 Luna (xhigh): score by benchmark
Codex – GPT-6 Luna (xhigh): by benchmark
BenchmarkScoreTasks with a decided attemptCost per taskTime per task
GeoAgentBench9950 of 50$0.00502.0 min
GeoBenchX60173 of 173$0.0113.9 min
Earth-Bench5548 of 48$0.00783.3 min
GeoAnalystBench9519 of 19$0.00712.7 min

Score by task format and kind of work

Codex – GPT-6 Luna (xhigh): score by task format

Average pass@1 over the tasks of each task format, 0 to 100 · Higher is better

Codex – GPT-6 Luna (xhigh): score by task formatMulti-step analysis: 77 (partial coverage); Single-answer questions: 56 (partial coverage); Recognising infeasible requests: 62 (partial coverage)025507510077Multi-stepanalysis56Single-answerquestions62Recognisinginfeasiblerequests
How to read this chart

What this shows. The same attempts as the index, grouped by task format. Groups with few tasks are less certain. About these breakdowns.

Data table: Codex – GPT-6 Luna (xhigh): score by task format
Codex – GPT-6 Luna (xhigh): by task format
Task formatScoreTasks with a decided attempt
Multi-step analysis77151 of 151
Single-answer questions5660 of 60
Recognising infeasible requests6279 of 79

Codex – GPT-6 Luna (xhigh): score by kind of work

Average pass@1 over the tasks of each kind of work, 0 to 100 · Higher is better

Codex – GPT-6 Luna (xhigh): score by kind of workRaster and remote sensing: 75 (partial coverage); Vector and overlay: 78 (partial coverage); Networks and routing: 100 (partial coverage); Climate and time series: 75 (partial coverage); Spatial statistics and interpolation: 63 (partial coverage); Mapping and cartography: 64 (partial coverage); Recognising infeasible requests: 62 (partial coverage)025507510075Raster andremotesensing78Vector andoverlay100Networks androuting75Climate andtime series63Spatialstatisticsand64Mapping andcartography62Recognisinginfeasiblerequests
How to read this chart

What this shows. The same attempts as the index, grouped by kind of work. Groups with few tasks are less certain. About these breakdowns.

Data table: Codex – GPT-6 Luna (xhigh): score by kind of work
Codex – GPT-6 Luna (xhigh): by kind of work
Kind of workScoreTasks with a decided attempt
Raster and remote sensing7575 of 75
Vector and overlay7818 of 18
Networks and routing1005 of 5
Climate and time series7524 of 24
Spatial statistics and interpolation6324 of 24
Mapping and cartography6465 of 65
Recognising infeasible requests6279 of 79

Compared with other models

Geospatial Agent Index

Equal-weight average of the evaluation scores, 0 to 100 · Higher is better

  • Partial coverage
Geospatial Agent IndexClaude Code – Claude Opus 5.5 (high): 83* (partial coverage); Claude Code – Claude Opus 5.5 (xhigh): 82; Claude Code – Claude Opus 5.5 (medium): 81* (partial coverage); Claude Code – Claude Sonnet 5.5 (high): 81* (partial coverage); Codex – GPT-6.1 Sol (high): 80* (partial coverage); Codex – GPT-6.1 Sol (medium): 79* (partial coverage); Codex – GPT-6.1 Sol (xhigh): 78* (partial coverage); Claude Code – Claude Haiku 5.5 (xhigh): 78; Codex – GPT-6 Luna (xhigh): 77* (partial coverage); Claude Code – Claude Haiku 5.5 (medium): 77* (partial coverage); Claude Code – Claude Sonnet 5.5 (medium): 77* (partial coverage); Claude Code – Claude Haiku 5.5 (high): 75* (partial coverage); Terminus-2 – GLM-5.2 (max): 75* (partial coverage); Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high): 74* (partial coverage); Terminus-2 – DeepSeek V4 Flash (max): 74* (partial coverage); Terminus-2 – DeepSeek V4 Pro (high): 74; Claude Code – Claude Sonnet 5.5 (xhigh): 74; Terminus-2 – Kimi K2.7 Code (always on): 74; Terminus-2 – Kimi K2.6 (always on): 72; Terminus-2 – GLM-5.2 (high): 71* (partial coverage); Terminus-2 – GLM-5.3 (low): 71* (partial coverage); Terminus-2 – DeepSeek V4 Flash (high): 70* (partial coverage); Terminus-2 – GLM-5.2 (none): 69* (partial coverage); Terminus-2 – GLM-5.3 (max, default): 67* (partial coverage); Terminus-2 – GLM-5.3-Flash (low): 66* (partial coverage); Terminus-2 – Gemma 4 26B A4B (thinking on): 62; Terminus-2 – GLM-5.3-Flash (max): 56* (partial coverage); Terminus-2 – gpt-oss-120b (high): 52; Terminus-2 – GLM-4.7-Flash (thinking on): 40; Terminus-2 – gpt-oss-120b (low): 37; Terminus-2 – Qwen 3.8 27B (xhigh): 62* (partial coverage)025507510083*Claude CodeClaude Opus5.5 (high)82Claude CodeClaude Opus5.5 (xhigh)81*Claude CodeClaude Opus5.5 (medium)81*Claude CodeClaude Sonnet5.5 (high)80*CodexGPT-6.1 Sol(high)79*CodexGPT-6.1 Sol(medium)78*CodexGPT-6.1 Sol(xhigh)78Claude CodeClaude Haiku5.5 (xhigh)77*CodexGPT-6 Luna(xhigh)77*Claude CodeClaude Haiku5.5 (medium)77*Claude CodeClaude Sonnet5.5 (medium)75*Claude CodeClaude Haiku5.5 (high)75*Terminus-2GLM-5.2 (max)74*Terminus-2Claude Sonnet4.6(adaptive,74*Terminus-2DeepSeek V4Flash (max)74Terminus-2DeepSeek V4Pro (high)74Claude CodeClaude Sonnet5.5 (xhigh)74Terminus-2Kimi K2.7Code (alwayson)72Terminus-2Kimi K2.6(always on)71*Terminus-2GLM-5.2(high)71*Terminus-2GLM-5.3 (low)70*Terminus-2DeepSeek V4Flash (high)69*Terminus-2GLM-5.2(none)67*Terminus-2GLM-5.3 (max,default)66*Terminus-2GLM-5.3-Flash(low)62Terminus-2Gemma 4 26BA4B (thinkingon)56*Terminus-2GLM-5.3-Flash(max)52Terminus-2gpt-oss-120b(high)40Terminus-2GLM-4.7-Flash(thinking on)37Terminus-2gpt-oss-120b(low)62*Terminus-2Qwen 3.8 27B(xhigh)
How to read this chart

What this metric means. The equal-weight average of the 4 evaluation scores. Each evaluation score is the average pass@1 over that benchmark's tasks, from 0 to 100.

Not yet ranked. Entries without a score on every benchmark come after the ranked entries and are not ranked: their index averages only the benchmarks they have.

Partial coverage. A hatched bar, and an asterisk in tables, marks an entry that does not yet have results for every planned attempt. While attempts await review, the table gives the range the score could still reach.

Partial attempts. Each task is planned 3 times. An entry with fewer attempts per task is in the index with the marker "partial: n of 3 attempts", where n is its attempts per task; its score rests on fewer attempts and can change more.

Data table: Geospatial Agent Index
Geospatial Agent Index
ModelCreatorGeospatial Agent IndexRange while attempts are pendingBenchmarks coveredCoverage (attempts)
Claude Code – Claude Opus 5.5 (high)Anthropic83* (partial: 1.7 of 3 attempts)71 to 834 of 4422 of 870 planned attempts
Claude Code – Claude Opus 5.5 (xhigh)Anthropic82 (partial: 1 of 3 attempts)No pending judgements4 of 4290 of 290 planned attempts
Claude Code – Claude Opus 5.5 (medium)Anthropic81* (partial: 1.8 of 3 attempts)70 to 814 of 4461 of 870 planned attempts
Claude Code – Claude Sonnet 5.5 (high)Anthropic81* (partial: 1.9 of 3 attempts)71 to 824 of 4479 of 870 planned attempts (236 excluded)
Codex – GPT-6.1 Sol (high)OpenAI80* (partial: 2.8 of 3 attempts)72 to 804 of 4753 of 870 planned attempts
Codex – GPT-6.1 Sol (medium)OpenAI79* (partial: 2.6 of 3 attempts)No pending judgements; coverage incomplete4 of 4749 of 870 planned attempts (6 excluded)
Codex – GPT-6.1 Sol (xhigh)OpenAI78* (partial: 2 of 3 attempts)74 to 794 of 4551 of 870 planned attempts
Claude Code – Claude Haiku 5.5 (xhigh)Anthropic78 (partial: 1 of 3 attempts)No pending judgements4 of 4290 of 290 planned attempts
Codex – GPT-6 Luna (xhigh)OpenAI77* (partial: 1.9 of 3 attempts)72 to 784 of 4512 of 870 planned attempts
Claude Code – Claude Haiku 5.5 (medium)Anthropic77* (partial: 1.8 of 3 attempts)66 to 784 of 4462 of 870 planned attempts
Claude Code – Claude Sonnet 5.5 (medium)Anthropic77* (partial: 2 of 3 attempts)67 to 774 of 4500 of 870 planned attempts (229 excluded)
Claude Code – Claude Haiku 5.5 (high)Anthropic75* (partial: 1.7 of 3 attempts)66 to 774 of 4422 of 870 planned attempts
Terminus-2 – GLM-5.2 (max)Z.ai75* (partial: 1.1 of 3 attempts)72 to 754 of 4306 of 870 planned attempts
Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high)Anthropic74* (partial: 1.1 of 3 attempts)73 to 754 of 4306 of 870 planned attempts (133 excluded)
Terminus-2 – DeepSeek V4 Flash (max)DeepSeek74* (partial: 1.1 of 3 attempts)73 to 744 of 4321 of 870 planned attempts
Terminus-2 – DeepSeek V4 Pro (high)DeepSeek74No pending judgements4 of 4870 of 870 planned attempts
Claude Code – Claude Sonnet 5.5 (xhigh)Anthropic74 (partial: 1 of 3 attempts)No pending judgements4 of 4290 of 290 planned attempts
Terminus-2 – Kimi K2.7 Code (always on)Moonshot AI74No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – Kimi K2.6 (always on)Moonshot AI72No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – GLM-5.2 (high)Z.ai71* (partial: 1 of 3 attempts)65 to 724 of 4273 of 290 planned attempts
Terminus-2 – GLM-5.3 (low)Z.ai71* (partial: 0.9 of 3 attempts)61 to 724 of 4238 of 290 planned attempts
Terminus-2 – DeepSeek V4 Flash (high)DeepSeek70*70 to 714 of 4860 of 870 planned attempts
Terminus-2 – GLM-5.2 (none)Z.ai69* (partial: 1.4 of 3 attempts)62 to 724 of 4356 of 870 planned attempts
Terminus-2 – GLM-5.3 (max, default)Z.ai67* (partial: 1.1 of 3 attempts)66 to 684 of 4310 of 870 planned attempts
Terminus-2 – GLM-5.3-Flash (low)Z.ai66* (partial: 1.1 of 3 attempts)59 to 674 of 4288 of 870 planned attempts (3 excluded)
Terminus-2 – Gemma 4 26B A4B (thinking on)Google62No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – GLM-5.3-Flash (max)Z.ai56* (partial: 1.1 of 3 attempts)54 to 564 of 4298 of 870 planned attempts
Terminus-2 – gpt-oss-120b (high)OpenAI52No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – GLM-4.7-Flash (thinking on)Z.ai40No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – gpt-oss-120b (low)OpenAI37No pending judgements4 of 4870 of 870 planned attempts
Terminus-2 – Qwen 3.8 27B (xhigh)Alibaba (Qwen)62* (partial: 0.3 of 3 attempts)No pending judgements; coverage incomplete3 of 4101 of 870 planned attempts

Geospatial Agent Index vs. cost per task

Higher and further left is better

  • Partial coverage
  • Most attractive quadrant
  • Pareto line
Geospatial Agent Index vs. cost per taskTerminus-2 – Claude Sonnet 4.6 (adaptive, effort high): Geospatial Agent Index 74.42, cost per task (usd) $0.278; Terminus-2 – DeepSeek V4 Flash (max): Geospatial Agent Index 74.25, cost per task (usd) $0.067; Terminus-2 – DeepSeek V4 Flash (high): Geospatial Agent Index 69.81, cost per task (usd) $0.067; Terminus-2 – DeepSeek V4 Pro (high): Geospatial Agent Index 74.2, cost per task (usd) $0.273; Terminus-2 – Gemma 4 26B A4B (thinking on): Geospatial Agent Index 61.62, cost per task (usd) $0.016; Terminus-2 – GLM-4.7-Flash (thinking on): Geospatial Agent Index 39.77, cost per task (usd) $0.030; Terminus-2 – GLM-5.2 (high): Geospatial Agent Index 71.12, cost per task (usd) $0.239; Terminus-2 – GLM-5.2 (none): Geospatial Agent Index 68.79, cost per task (usd) $0.103; Terminus-2 – GLM-5.2 (max): Geospatial Agent Index 74.83, cost per task (usd) $0.362; Terminus-2 – GLM-5.3-Flash (low): Geospatial Agent Index 65.51, cost per task (usd) $0.028; Terminus-2 – GLM-5.3-Flash (max): Geospatial Agent Index 55.62, cost per task (usd) $0.042; Terminus-2 – GLM-5.3 (low): Geospatial Agent Index 70.99, cost per task (usd) $0.250; Terminus-2 – GLM-5.3 (max, default): Geospatial Agent Index 67.44, cost per task (usd) $0.363; Terminus-2 – gpt-oss-120b (low): Geospatial Agent Index 37.05, cost per task (usd) $0.0088; Terminus-2 – gpt-oss-120b (high): Geospatial Agent Index 52.1, cost per task (usd) $0.026; Terminus-2 – Kimi K2.6 (always on): Geospatial Agent Index 72.4, cost per task (usd) $0.072; Terminus-2 – Kimi K2.7 Code (always on): Geospatial Agent Index 73.7, cost per task (usd) $0.119; Terminus-2 – Qwen 3.8 27B (xhigh): Geospatial Agent Index 62.5, cost per task (usd) $0.066; Claude Code – Claude Opus 5.5 (medium): Geospatial Agent Index 80.86, cost per task (usd) $0.138; Claude Code – Claude Opus 5.5 (high): Geospatial Agent Index 82.76, cost per task (usd) $0.175; Claude Code – Claude Sonnet 5.5 (medium): Geospatial Agent Index 76.59, cost per task (usd) $0.052; Claude Code – Claude Sonnet 5.5 (high): Geospatial Agent Index 80.76, cost per task (usd) $0.062; Claude Code – Claude Haiku 5.5 (medium): Geospatial Agent Index 76.6, cost per task (usd) $0.0075; Claude Code – Claude Haiku 5.5 (high): Geospatial Agent Index 74.94, cost per task (usd) $0.010; Claude Code – Claude Opus 5.5 (xhigh): Geospatial Agent Index 82.32, cost per task (usd) $0.375; Claude Code – Claude Sonnet 5.5 (xhigh): Geospatial Agent Index 74.18, cost per task (usd) $0.163; Claude Code – Claude Haiku 5.5 (xhigh): Geospatial Agent Index 77.89, cost per task (usd) $0.023; Codex – GPT-6.1 Sol (medium): Geospatial Agent Index 79.48, cost per task (usd) $0.060; Codex – GPT-6.1 Sol (high): Geospatial Agent Index 79.77, cost per task (usd) $0.080; Codex – GPT-6.1 Sol (xhigh): Geospatial Agent Index 78.07, cost per task (usd) $0.124; Codex – GPT-6 Luna (xhigh): Geospatial Agent Index 77.44, cost per task (usd) $0.0092Most attractive quadrant0255075100$0.0010$0.0020$0.0050$0.010$0.020$0.050$0.100$0.200$0.500$1.00Cost per task (USD) (log scale)Geospatial Agent IndexClaude Code – Claude Opus 5.5 (high)Claude Code – Claude Opus 5.5 (xhigh)Claude Code – Claude Opus 5.5 (medium)Claude Code – Claude Sonnet 5.5 (high)Codex – GPT-6.1 Sol (high)Codex – GPT-6.1 Sol (medium)Codex – GPT-6.1 Sol (xhigh)Claude Code – Claude Haiku 5.5 (xhigh)Codex – GPT-6 Luna (xhigh)Claude Code – Claude Haiku 5.5 (medium)Claude Code – Claude Sonnet 5.5 (medium)Claude Code – Claude Haiku 5.5 (high)Terminus-2 – GLM-5.2 (max)Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high)Terminus-2 – DeepSeek V4 Flash (max)Terminus-2 – DeepSeek V4 Pro (high)Claude Code – Claude Sonnet 5.5 (xhigh)Terminus-2 – Kimi K2.7 Code (always on)Terminus-2 – Kimi K2.6 (always on)Terminus-2 – GLM-5.2 (high)Terminus-2 – GLM-5.3 (low)Terminus-2 – DeepSeek V4 Flash (high)Terminus-2 – GLM-5.2 (none)Terminus-2 – GLM-5.3 (max, default)Terminus-2 – GLM-5.3-Flash (low)Terminus-2 – Qwen 3.8 27B (xhigh)Terminus-2 – Gemma 4 26B A4B (thinking on)Terminus-2 – GLM-5.3-Flash (max)Terminus-2 – gpt-oss-120b (high)Terminus-2 – GLM-4.7-Flash (thinking on)Terminus-2 – gpt-oss-120b (low)
Data table: Geospatial Agent Index vs. cost per task
Geospatial Agent Index vs. cost per task
ModelGeospatial Agent IndexCost per task (USD)On the Pareto line
Terminus-2 – Claude Sonnet 4.6 (adaptive, effort high)74* (partial: 1.1 of 3 attempts)$0.278No
Terminus-2 – DeepSeek V4 Flash (max)74* (partial: 1.1 of 3 attempts)$0.067No
Terminus-2 – DeepSeek V4 Flash (high)70*$0.067No
Terminus-2 – DeepSeek V4 Pro (high)74$0.273No
Terminus-2 – Gemma 4 26B A4B (thinking on)62$0.016No
Terminus-2 – GLM-4.7-Flash (thinking on)40$0.030No
Terminus-2 – GLM-5.2 (high)71* (partial: 1 of 3 attempts)$0.239No
Terminus-2 – GLM-5.2 (none)69* (partial: 1.4 of 3 attempts)$0.103No
Terminus-2 – GLM-5.2 (max)75* (partial: 1.1 of 3 attempts)$0.362No
Terminus-2 – GLM-5.3-Flash (low)66* (partial: 1.1 of 3 attempts)$0.028No
Terminus-2 – GLM-5.3-Flash (max)56* (partial: 1.1 of 3 attempts)$0.042No
Terminus-2 – GLM-5.3 (low)71* (partial: 0.9 of 3 attempts)$0.250No
Terminus-2 – GLM-5.3 (max, default)67* (partial: 1.1 of 3 attempts)$0.363No
Terminus-2 – gpt-oss-120b (low)37$0.0088No
Terminus-2 – gpt-oss-120b (high)52$0.026No
Terminus-2 – Kimi K2.6 (always on)72$0.072No
Terminus-2 – Kimi K2.7 Code (always on)74$0.119No
Terminus-2 – Qwen 3.8 27B (xhigh)62* (partial: 0.3 of 3 attempts)$0.066No
Claude Code – Claude Opus 5.5 (medium)81* (partial: 1.8 of 3 attempts)$0.138Yes
Claude Code – Claude Opus 5.5 (high)83* (partial: 1.7 of 3 attempts)$0.175Yes
Claude Code – Claude Sonnet 5.5 (medium)77* (partial: 2 of 3 attempts)$0.052No
Claude Code – Claude Sonnet 5.5 (high)81* (partial: 1.9 of 3 attempts)$0.062Yes
Claude Code – Claude Haiku 5.5 (medium)77* (partial: 1.8 of 3 attempts)$0.0075Yes
Claude Code – Claude Haiku 5.5 (high)75* (partial: 1.7 of 3 attempts)$0.010No
Claude Code – Claude Opus 5.5 (xhigh)82 (partial: 1 of 3 attempts)$0.375No
Claude Code – Claude Sonnet 5.5 (xhigh)74 (partial: 1 of 3 attempts)$0.163No
Claude Code – Claude Haiku 5.5 (xhigh)78 (partial: 1 of 3 attempts)$0.023Yes
Codex – GPT-6.1 Sol (medium)79* (partial: 2.6 of 3 attempts)$0.060Yes
Codex – GPT-6.1 Sol (high)80* (partial: 2.8 of 3 attempts)$0.080No
Codex – GPT-6.1 Sol (xhigh)78* (partial: 2 of 3 attempts)$0.124No
Codex – GPT-6 Luna (xhigh)77* (partial: 1.9 of 3 attempts)$0.0092Yes

Head-to-head comparisons

Specification and settings

Developer
OpenAI
Context window
No data
Image input
No data
Reasoning setting
xhigh
Temperature
Provider default
Maximum output tokens
No data
Input price per 1M tokens
$0.10
Cached input price per 1M tokens
$0.010
Output price per 1M tokens
$0.50

Why this model was chosen: model selection.