Geospatial agent benchmark · not-downloaded

GeoMMBench / GeoMMAgent

Also known as: GeoMMBench, GeoMMAgent, GeoMM-AGI

Evaluates: model answers only

Expert multimodal geoscience and remote-sensing MCQ accuracy across six sensor modalities and four geoscience disciplines. CVPR 2026 Highlight.

At a glance

1053Tasks
NoAgents run real code
YesData + code public
YesChecked by recomputation
CC BY 4.0Data licence

Sources

every claim on this page traces to these records

In shortAgents do not run code inside this benchmark, and answers are checked by recomputation, not by another model's opinion. Dataset and code are published, so you can verify claims yourself. Closed-form MCQ does not prove execution. A model scoring 90% may still fail to produce a cloud-free NDVI composite or execute a spatial workflow.

Evidence boundaries

what a score on this benchmark does and does not tell you

A good score shows

Expert multimodal geoscience and remote-sensing MCQ accuracy across six sensor modalities and four geoscience disciplines. CVPR 2026 Highlight.

A good score does not show

Closed-form MCQ does not prove execution. A model scoring 90% may still fail to produce a cloud-free NDVI composite or execute a spatial workflow.

Arena status

not-downloaded — Not downloaded. VLM sidecar benchmark, multimodal model-selection comparator, and possible small pinned image-understanding seed.

All recorded facts (22 fields)
Evaluates
model
Task families
scene classification,object detection,change detection,spectral analysis,spatial reasoning,terrain characterisation,GNSS pseudorange calculations
Task count
1053
Difficulty
expert-level (derived from geoscience curricula, not crowd-sourced)
Geography
varied (remote sensing, photogrammetry, GIS, GNSS disciplines)
Modality
mixed
Input types
image,text (question),multiple-choice options A-D
Output types
multiple-choice answer (A/B/C/D)
Execution environment
VLM inference (zero-shot, 36+ VLMs tested); GeoMMAgent runs YOLO11 and DeepLabV3+ perception models locally
Ground truth method
expert-derived multiple-choice questions with verified answers
Verifier method
MCQ accuracy; optional self-evaluation dimensions (logic, spatial reasoning, domain validity, accuracy)
Deterministic verification
yes
Scoring dimensions
MCQ accuracy,optional self-eval: logic, spatial reasoning, domain validity, accuracy
Aggregation formula
MCQ accuracy
Trials policy
36+ VLMs tested under zero-shot conditions
Variance reporting
pending verification (per-model accuracy numbers not published in public README at time of analysis)
Data licence
CC BY 4.0 (dataset)
Code licence
Apache 2.0 (GeoMMAgent framework)
Contamination concerns
public Hugging Face dataset; 37-question public validation split is small and may yield high variance in external reproductions
Reproducibility status
medium (dataset public on HuggingFace; per-model accuracy in paper PDF; arXiv identifier confirmed as 2604.08896)
First published
2026-04
Latest known update
2026-05-02

Machine-readable record

This entry is available as JSON at /api/registry#geommbench-geommagent. See the registry endpoint.