How can an agent keep spatial meaning intact?
A source can describe an observation, a mapped entity, a prediction or a decision. We are interested in systems that retain that difference rather than flattening it into generic context.
Axis Spatial Research
Axis Spatial is a research, engineering, and product company building spatial intelligence for AI agents: systems that can understand context, investigate through real software, and make their work available for review.
Current questions
A source can describe an observation, a mapped entity, a prediction or a decision. We are interested in systems that retain that difference rather than flattening it into generic context.
Spatial work often turns on one calculation, image, record or rule. We study how an agent can make that next action legible and connected to the evidence behind it.
A program running is different from a result being spatially sound. Geometry, coordinate systems, units, extent, time and provenance need their own checks.
Nearby records can describe one entity, adjacent entities, different times or competing explanations. Identity has to remain testable while evidence is incomplete.
The evidence, unresolved alternatives, tool state and next accepted action must outlive an interruption without relying on a chat transcript.
Model, provider, prompt, harness, tools, runtime, retry policy, data and verifier all shape the result. A model name alone is not a valid comparison unit.
Public evidence layer
Axis Spatial Arena is the public evidence layer for spatial-agent work. The comparison unit is the complete system: model, provider, prompt, harness, tools, runtime, retry path, benchmark, dataset and verifier.
Method
An evidence record should say what was configured, what happened, which artefacts can be inspected, and what the record does not prove. That is how an observed result remains useful without becoming a general claim.
Featured evidence record · 27 July 2026
A one-task benchmark audit, not a suite result. GeoAnalystBench Task 34 showed that two observed model routes could generate code, produce inspectable artefacts and return the same benchmark-compatible result.
Verdict one · benchmark match
Kimi K2.6 and Kimi K2.7 Code each produced 84.3348706%, compatible with the published benchmark outcome.
Verdict two · spatial validity
A benchmark match does not establish that the result is spatially valid for the stated question. Task 34 does not yet have an accepted spatial-validity gold.
Inspectable artefacts
Kimi K2.6
Kimi K2.7 Code
The source inputs are not served. The public record contains generated code, result JSON, maps and the evidence manifest. The source benchmark remains the relevant upstream reference: GeoAnalystBench source.
Limitations
01No rank, pass-rate or reliability claim.
02No model comparison, reproduction or full-suite claim.
03No production Worker/E2B claim.
04No accepted spatial-validity gold.
05Eight lanes are single routed observations, not reliability estimates.
06One task is an operating-path proof, not a general result.
07Arena is maintained by Axis Spatial, not an independent certification body.
Insights
Read our current thinking on spatial intelligence, evidence and applied systems.
Visit InsightsOpen Lab
When reference workflows, task packs, spatial verifiers or fixtures are ready to inspect and run, they will appear here with their limits and licence. We will not present placeholders as if they are available.