Skip to main content

Research and evidence

We test what spatial AI can and cannot do.

Axis Spatial researches and tests AI agent systems for geospatial work.

We publish the result, the failure, the evidence, and the limit.

Current questions

The problems behind reliable spatial agents.

01

How can an agent keep spatial meaning intact?

A source can describe an observation, a mapped entity, a prediction or a decision. We are interested in systems that retain that difference rather than flattening it into generic context.

02

What should an agent investigate next?

Spatial work often turns on one calculation, image, record or rule. We study how an agent can make that next action legible and connected to the evidence behind it.

03

What makes a spatial outcome acceptable?

A program running is different from a result being spatially sound. Geometry, coordinate systems, units, extent, time and provenance need their own checks.

04

When are two observations the same physical thing?

Nearby records can describe one entity, adjacent entities, different times or competing explanations. Identity has to remain testable while evidence is incomplete.

05

What must survive a long-running investigation?

The evidence, unresolved alternatives, tool state and next accepted action must outlive an interruption without relying on a chat transcript.

06

How should complete agent systems be compared?

Model, provider, prompt, harness, tools, runtime, retry policy, data and verifier all shape the result. A model name alone is not a valid comparison unit.

Axis Spatial Arena

A public index of geospatial AI benchmarks.

Arena records what each benchmark tests, how it checks a result, and what the result cannot show. The live index links public benchmark entries to their sources and methods.

It also records selected Axis runs with the task setup, outputs, checks, cost, time, failures, and known limits.

Spatial tasks connected to output files, checks, and a reviewable evidence record.
01

Public benchmark entries

The index links each entry to its available paper, data, code, and reported method.

02

One comparison format

Each entry states what was tested, how results were checked, and where evidence is missing.

03

Clear limits

The index does not certify a system or turn one result into a product claim.

Current evidence boundary

Read each benchmark with its setup and limits.

The live registry records the source, execution setup, checks, and limits that are available for each entry. A benchmark score is not a product claim or proof that a spatial output is correct.

Open the current benchmark registry (opens in a new tab)

What this record does not prove

Clear limits.

No rank, pass-rate or reliability claim.

No model comparison, reproduction or full-suite claim.

No accepted spatial-validity gold for the task.

A benchmark entry is an evidence record, not a general product result.

Arena is maintained by Axis Spatial, not an independent certification body.

Practical library

Guides and blog.

These earlier guides cover agents, cloud platforms, automation, and workflow migration. Their URLs stay intact.

AI Agents for Geospatial Work: Beyond the Chatbot

What geospatial agents need beyond a chat interface: context, tools, checks, evidence, and clear limits.

Read guide (opens in a new tab)

AI Agents in GIS: Beyond the Hype

A practical account of where AI agents help GIS teams and where human review still matters.

Read guide (opens in a new tab)

Geospatial Workflow Automation

A guide to workflow choices, implementation stages, and the work that should not be automated.

Read guide (opens in a new tab)

Cloud-Native Geospatial Formats

How formats such as COG, GeoParquet, and STAC change the design of spatial workflows.

Read guide (opens in a new tab)

Start here

Do you need to test a geospatial agent?

We can define a task, run the system, inspect the output files, and state what the result does and does not show.