Source paper
the peer-reviewed or arXiv publication that defines the benchmark
Geospatial agent benchmark · none
Also known as: UrbanSARFloods, Sentinel-1 Flood Benchmark
Evaluates: model answers only
For SAR flood detection, the best baselines score 0.52-0.65 F1 on urban chips versus 0.78-0.85 on open-area chips. Typical scenes contain under 5% flooded pixels; class imbalance is the main limiting factor.
the peer-reviewed or arXiv publication that defines the benchmark
In shortAgents do not run code inside this benchmark, and answers are checked by recomputation, not by another model's opinion. The underlying data and code are not published at a verified link. Agent execution, workflow generation, or production deployment. Does not test code generation or tool use.
For SAR flood detection, the best baselines score 0.52-0.65 F1 on urban chips versus 0.78-0.85 on open-area chips. Typical scenes contain under 5% flooded pixels; class imbalance is the main limiting factor.
Agent execution, workflow generation, or production deployment. Does not test code generation or tool use.
none — Not downloaded. Related EO benchmark family. Used as a reference standard in the Bearlin v1 benchmark spec for SAR flood task acceptance criteria (flooded km2 within 25% of UrbanSARFloods reference).
This entry is available as JSON at /api/registry#urbansarfloods. See the registry endpoint.