Geospatial agent benchmark · none

RSRCC

Also known as: RSRCC, Remote Sensing Regional Change Comprehension

Evaluates: model answers only

Fine-grained semantic reasoning about localised changes in remote sensing image pairs. Top models achieve approximately 60-65% on fine-grained questions versus 80%+ on global binary questions. The difficulty gap is largest for spectrally similar changes.

At a glance

126000Tasks
NoAgents run real code
NoData + code public
YesChecked by recomputation
Data licence

Sources

every claim on this page traces to these records

In shortAgents do not run code inside this benchmark, and answers are checked by recomputation, not by another model's opinion. The underlying data and code are not published at a verified link. Code execution, workflow generation, or artifact production. Does not test agent tool-use or production deployment.

Evidence boundaries

what a score on this benchmark does and does not tell you

A good score shows

Fine-grained semantic reasoning about localised changes in remote sensing image pairs. Top models achieve approximately 60-65% on fine-grained questions versus 80%+ on global binary questions. The difficulty gap is largest for spectrally similar changes.

A good score does not show

Code execution, workflow generation, or artifact production. Does not test agent tool-use or production deployment.

Arena status

none — Not downloaded. Remote-sensing change-comprehension benchmark. The annotation pipeline may be adaptable for generating training data for fine-tuned summarisation models.

All recorded facts (22 fields)
Evaluates
model
Task families
remote sensing change comprehension,fine-grained spatial localisation,bi-temporal image pair reasoning
Task count
126000
Difficulty
mixed (global binary change detection through fine-grained sub-region identification)
Geography
varied (bi-temporal remote sensing image pairs)
Modality
optical
Input types
bi-temporal remote sensing image pairs,natural language questions
Output types
fine-grained change descriptions with spatial localisation
Execution environment
MLLM or VLM inference
Ground truth method
hierarchical semi-supervised curation pipeline with Best-of-N ranking as final ambiguity-resolution stage
Verifier method
accuracy on fine-grained questions versus global binary questions
Deterministic verification
yes
Scoring dimensions
fine-grained change identification accuracy,global binary change detection accuracy
Aggregation formula
accuracy; difficulty gap measured between fine-grained and global binary questions
Trials policy
pending verification
Variance reporting
pending verification
Data licence
pending verification
Code licence
pending verification
Contamination concerns
large-scale public benchmark; contamination possible if MLLMs trained on it
Reproducibility status
medium (methodology public; dataset and code availability pending verification)
First published
2026-04
Latest known update
2026-05-02

Machine-readable record

This entry is available as JSON at /api/registry#rsrcc. See the registry endpoint.