Geospatial agent benchmark · not-downloaded

GEOBench-VLM

Also known as: GEOBench-VLM, geo-bench-vlm

Evaluates: model answers only

Geospatial VLM multiple-choice question-answering accuracy on remote-sensing imagery.

At a glance

pending verificationTasks
NoAgents run real code
PartlyData + code public
YesChecked by recomputation
Data licence

Sources

every claim on this page traces to these records

In shortAgents do not run code inside this benchmark, and answers are checked by recomputation, not by another model's opinion. Part of the evidence base (dataset or code, not both) is published. Not a primary spatial verifier. MCQ accuracy does not prove code execution, artifact production, or workflow correctness.

Evidence boundaries

what a score on this benchmark does and does not tell you

A good score shows

Geospatial VLM multiple-choice question-answering accuracy on remote-sensing imagery.

A good score does not show

Not a primary spatial verifier. MCQ accuracy does not prove code execution, artifact production, or workflow correctness.

Arena status

not-downloaded — Not downloaded. VLM sidecar benchmark and possible small pinned image-understanding seed.

All recorded facts (22 fields)
Evaluates
model
Task families
geospatial VLM multiple-choice question answering
Task count
pending verification
Difficulty
pending verification
Geography
pending verification
Modality
mixed
Input types
remote-sensing imagery,language instructions
Output types
multiple-choice answers
Execution environment
VLM inference (repo includes inference scripts)
Ground truth method
manually verified instructions and answer options
Verifier method
MCQ accuracy
Deterministic verification
yes
Scoring dimensions
MCQ accuracy
Aggregation formula
MCQ accuracy
Trials policy
pending verification
Variance reporting
pending verification
Data licence
pending verification
Code licence
pending verification
Contamination concerns
public Hugging Face dataset; contamination possible if VLMs trained on it
Reproducibility status
medium (Hugging Face dataset with inference scripts; GPU, model, and provider posture must be explicit)
First published
pending verification
Latest known update
pending verification

Machine-readable record

This entry is available as JSON at /api/registry#geobench-vlm. See the registry endpoint.