Geospatial agent benchmark · none

GeoAI Agency Primitives

Also known as: GeoAI Agency Primitives, Agency Primitives Benchmark

Evaluates: end-to-end agents

The paper proposes a nine-primitive framework and benchmark protocol for GeoAI agency. Implementation and validation remain future work.

At a glance

pending verificationTasks
YesAgents run real code
NoData + code public
YesChecked by recomputation
Data licence

Sources

every claim on this page traces to these records

In shortAgents run real code inside this benchmark, and answers are checked by recomputation, not by another model's opinion. The underlying data and code are not published at a verified link. Measured productivity gains, code execution correctness, artifact production, or spatial verification.

Evidence boundaries

what a score on this benchmark does and does not tell you

A good score shows

The paper proposes a nine-primitive framework and benchmark protocol for GeoAI agency. Implementation and validation remain future work.

A good score does not show

Measured productivity gains, code execution correctness, artifact production, or spatial verification.

Arena status

none — Not downloaded. Spatial-reasoning and capability framework beyond MapQA. Nine primitive taxonomy used as a capability framework reference in the GeoAI Analysis product plan.

All recorded facts (22 fields)
Evaluates
agent
Task families
navigation,perception,geo-referenced memory,action planning,spatial reasoning,tool composition,multi-step coordination,uncertainty quantification,human-in-the-loop delegation
Task count
pending verification
Difficulty
mixed (eight task types)
Geography
pending verification
Modality
mixed
Input types
GIS workflow tasks,geospatial data
Output types
analyst productivity metrics,primitive implementation scores
Execution environment
GeoAI assistant with varying subsets of nine agency primitives
Ground truth method
analyst productivity baseline comparison across eight task types
Verifier method
productivity gain measurement (percentage improvement versus unassisted baseline)
Deterministic verification
yes
Scoring dimensions
analyst productivity,per-primitive contribution (action planning, tool composition, geo-referenced memory, perception, spatial reasoning, uncertainty quantification)
Aggregation formula
productivity gain percentage versus unassisted baseline; per-primitive ablation
Trials policy
pending verification
Variance reporting
pending verification
Data licence
pending verification
Code licence
pending verification
Contamination concerns
pending verification
Reproducibility status
pending verification
First published
2026-04
Latest known update
pending verification

Machine-readable record

This entry is available as JSON at /api/registry#geoai-agency-primitives. See the registry endpoint.