Axis Spatial

Model selection

Which model configurations Geospatial Agent Benchmarks tests, why, and with which settings. As of 9 October 2026.

Selection principles

  • One provider and one API for every model. All models are open-weight models served by the same provider through the same API, so they share request handling, logging and per-request cost evidence, and differences come from the models rather than from provider plumbing.
  • Open-weight models in this round. Proprietary frontier models are not in this round; adding them is a later option.
  • A spread of developers, sizes and prices, so the results say something about cost against quality.
  • Matching Artificial Analysis identities where possible. Four of the models were chosen because each has an Artificial Analysis entry.
  • Each configuration is tested before use. Every model was probed with the exact settings it runs with (reasoning on, explicit output limit), and the probe is recorded.

The models

Each model is named with its reasoning level in brackets, as Artificial Analysis does: the effort level where the model has one (high, xhigh), (Reasoning) where reasoning is on without levels, and (Non-reasoning) where it is off. Each reasoning level is a separate entry with its own results and page.

  1. Kimi K2.7 Code (Reasoning) (Moonshot AI). Coding-oriented and the highest-priced of the first four ($0.95 / $4.00 per million input / output tokens; cached input $0.19). Image input: yes. Reasoning always on (no control exposed).
  2. GLM-4.7-Flash (Reasoning) (Z.ai). Economical text model ($0.06 / $0.40). Image input: no. Thinking switched on explicitly; a probe with it off gave a wrong answer.
  3. Gemma 4 26B A4B (Reasoning) (Google). Economical, image-capable ($0.10 / $0.30). Reasoning on by default.
  4. gpt-oss-120b (high) (OpenAI, open weights). Text-only reasoning model from a fourth developer ($0.35 / $0.75). Reasoning effort high (probe: low used 113 output tokens, high 516, on the same question). Its default output limit is only 256 tokens, so the limit is always set.
  5. DeepSeek V4 Pro (high) (DeepSeek). Larger DeepSeek model, 1M-token context ($1.32 / $3.96; cached input $0.044). Image input: no. Reasoning effort high.
  6. DeepSeek V4 Flash (high) (DeepSeek). Faster, cheaper sibling ($0.44 / $1.32; cached input $0.014). Image input: no. Reasoning effort high.
  7. Qwen 3.8 27B (xhigh) (Alibaba, Qwen). Mid-size, image-capable ($0.45 / $3.20; cached input $0.05). Reasoning effort xhigh, its default.

Prices are the serving provider's published list prices on the dates above; published costs use them.

Settings common to all models

  • Temperature 0.6 with reasoning on (Artificial Analysis's rule), unless a model's developer recommends otherwise.
  • Output limit 65,536 tokens, set explicitly; every model accepted it in a probe.
  • Reasoning passed back between agent turns.
  • Same agent (Terminus-2), same sandbox, same task inputs and limits for every model.

Considered and not used

  • Llama 4 Scout (Meta): not included in this round.
  • Older or overlapping versions (Kimi K2.6, gpt-oss-20b, Qwen3 30B, a DeepSeek R1 distillation): newer or larger siblings are already in the set.
  • GLM 5.2, 5.3 and 5.3 Flash: not in this round; candidates for the next one.
  • Proprietary frontier models (Claude, GPT, Gemini): outside this open-model comparison; possible later.