Geospatial agent benchmark · none

UrbanSARFloods

Also known as: UrbanSARFloods, Sentinel-1 Flood Benchmark

Evaluates: model answers only

For SAR flood detection, the best baselines score 0.52-0.65 F1 on urban chips versus 0.78-0.85 on open-area chips. Typical scenes contain under 5% flooded pixels; class imbalance is the main limiting factor.

At a glance

8879Tasks
NoAgents run real code
NoData + code public
YesChecked by recomputation
Data licence

Sources

every claim on this page traces to these records

In shortAgents do not run code inside this benchmark, and answers are checked by recomputation, not by another model's opinion. The underlying data and code are not published at a verified link. Agent execution, workflow generation, or production deployment. Does not test code generation or tool use.

Evidence boundaries

what a score on this benchmark does and does not tell you

A good score shows

For SAR flood detection, the best baselines score 0.52-0.65 F1 on urban chips versus 0.78-0.85 on open-area chips. Typical scenes contain under 5% flooded pixels; class imbalance is the main limiting factor.

A good score does not show

Agent execution, workflow generation, or production deployment. Does not test code generation or tool use.

Arena status

none — Not downloaded. Related EO benchmark family. Used as a reference standard in the Bearlin v1 benchmark spec for SAR flood task acceptance criteria (flooded km2 within 25% of UrbanSARFloods reference).

All recorded facts (22 fields)
Evaluates
model
Task families
SAR flood mapping,pixel-level flood or no-flood classification,land-cover-stratified evaluation
Task count
8879
Difficulty
mixed (class-imbalanced; urban flooding is hardest due to double-bounce SAR artefacts)
Geography
5 continents, 807,500 square kilometres
Modality
SAR
Input types
Sentinel-1 SLC chips
Output types
pixel-level flood or no-flood labels
Execution environment
deep learning model evaluation (pixel-level segmentation)
Ground truth method
pixel-level flood or no-flood labels stratified by land-cover class and continent
Verifier method
F1 score per chip, stratified by land-cover class and urban versus open-area
Deterministic verification
yes
Scoring dimensions
F1 score,urban chip F1,open-area chip F1,class-imbalance robustness
Aggregation formula
F1 score per chip; stratified by land-cover class and continent
Trials policy
pending verification (several deep learning baselines evaluated)
Variance reporting
pending verification
Data licence
pending verification
Code licence
pending verification
Contamination concerns
large-scale benchmark; contamination possible if flood detection models trained on it
Reproducibility status
medium (methodology public; dataset availability pending verification)
First published
2024-06
Latest known update
pending verification

Machine-readable record

This entry is available as JSON at /api/registry#urbansarfloods. See the registry endpoint.