Geospatial agent benchmark · not-downloaded

GEO-Bench-2

Also known as: GEO-Bench-2

Evaluates: model answers only

EO foundation-model downstream task quality across 19 datasets spanning multiple sensors.

At a glance

19Tasks
NoAgents run real code
PartlyData + code public
YesChecked by recomputation
varies by datasetData licence

Sources

every claim on this page traces to these records

In shortAgents do not run code inside this benchmark, and answers are checked by recomputation, not by another model's opinion. Part of the evidence base (dataset or code, not both) is published. Agent or tool-use execution. Does not test code generation, workflow planning, or artifact production.

Evidence boundaries

what a score on this benchmark does and does not tell you

A good score shows

EO foundation-model downstream task quality across 19 datasets spanning multiple sensors.

A good score does not show

Agent or tool-use execution. Does not test code generation, workflow planning, or artifact production.

Arena status

not-downloaded — Not downloaded. Planned as model-selection benchmark and EO or GFM comparator. Pin HF dataset versions, licences, selected split, and cache refs before use.

All recorded facts (22 fields)
Evaluates
model
Task families
EO foundation-model downstream tasks: classification, segmentation, change detection, regression, retrieval
Task count
19
Difficulty
varied by dataset
Geography
global (multi-sensor EO coverage)
Modality
mixed
Input types
raster EO,multi-sensor satellite imagery
Output types
classification labels,segmentation masks,regression values,change detection maps
Execution environment
model evaluation (foundation model downstream; not agent execution)
Ground truth method
labels and targets in official Hugging Face dataset versions
Verifier method
accuracy, F1, Jaccard or IoU, RMSE depending on dataset
Deterministic verification
yes
Scoring dimensions
accuracy,F1,Jaccard/IoU,RMSE
Aggregation formula
per-dataset metrics; no universal aggregate across 19 datasets
Trials policy
pending verification
Variance reporting
pending verification
Data licence
varies by dataset (Hugging Face)
Code licence
pending verification
Contamination concerns
public datasets; contamination possible if foundation models trained on them
Reproducibility status
high (Hugging Face datasets with pinned labels and targets)
First published
2025-11
Latest known update
pending verification

Machine-readable record

This entry is available as JSON at /api/registry#geo-bench-2. See the registry endpoint.