Source paper
the peer-reviewed or arXiv publication that defines the benchmark
Geospatial agent benchmark · none
Also known as: RSRCC, Remote Sensing Regional Change Comprehension
Evaluates: model answers only
Fine-grained semantic reasoning about localised changes in remote sensing image pairs. Top models achieve approximately 60-65% on fine-grained questions versus 80%+ on global binary questions. The difficulty gap is largest for spectrally similar changes.
the peer-reviewed or arXiv publication that defines the benchmark
In shortAgents do not run code inside this benchmark, and answers are checked by recomputation, not by another model's opinion. The underlying data and code are not published at a verified link. Code execution, workflow generation, or artifact production. Does not test agent tool-use or production deployment.
Fine-grained semantic reasoning about localised changes in remote sensing image pairs. Top models achieve approximately 60-65% on fine-grained questions versus 80%+ on global binary questions. The difficulty gap is largest for spectrally similar changes.
Code execution, workflow generation, or artifact production. Does not test agent tool-use or production deployment.
none — Not downloaded. Remote-sensing change-comprehension benchmark. The annotation pipeline may be adaptable for generating training data for fine-tuned summarisation models.
This entry is available as JSON at /api/registry#rsrcc. See the registry endpoint.