Source paper
the peer-reviewed or arXiv publication that defines the benchmark
Geospatial agent benchmark · not-downloaded
Also known as: GeoMMBench, GeoMMAgent, GeoMM-AGI
Evaluates: model answers only
Expert multimodal geoscience and remote-sensing MCQ accuracy across six sensor modalities and four geoscience disciplines. CVPR 2026 Highlight.
the peer-reviewed or arXiv publication that defines the benchmark
huggingface.co/datasets/AR-X/GeoMMBench
the published dataset the tasks are drawn from
github.com/Shihao-Cheng/GeoMMAgent
the official repository with harness, tasks and verifiers
In shortAgents do not run code inside this benchmark, and answers are checked by recomputation, not by another model's opinion. Dataset and code are published, so you can verify claims yourself. Closed-form MCQ does not prove execution. A model scoring 90% may still fail to produce a cloud-free NDVI composite or execute a spatial workflow.
Expert multimodal geoscience and remote-sensing MCQ accuracy across six sensor modalities and four geoscience disciplines. CVPR 2026 Highlight.
Closed-form MCQ does not prove execution. A model scoring 90% may still fail to produce a cloud-free NDVI composite or execute a spatial workflow.
not-downloaded — Not downloaded. VLM sidecar benchmark, multimodal model-selection comparator, and possible small pinned image-understanding seed.
This entry is available as JSON at /api/registry#geommbench-geommagent. See the registry endpoint.