Start with these reviewed benchmark records. Each explains the task, available software and data, checks, and limits of comparison.
Selected · Sources reviewed 24 Aug 2026
Executable GIS analysis with reference code and deterministic checks of vector, raster and tabular outputs.
Checks: L1 API F1, L2 reasoning similarity, L3 output verification (vector geometry diff via shapely, raster pixel diff via numpy, tabular value diff via pandas). Binary pass/fail.
Read the benchmark record
Selected · Sources reviewed 24 Aug 2026
Spatial reasoning questions grounded in OpenStreetMap, with a separate tool-assisted mode for routing and spatial queries.
Checks: accuracy (exact match against ground-truth answer)
Read the benchmark record
Selected · Sources reviewed 24 Aug 2026
Multi-turn environmental analysis through a defined tool API, checked with eight mechanistic tests per case.
Checks: eight mechanistic checks per case, including expected tools, required or forbidden text, numeric tolerance, chart generation, and loop limits
Read the benchmark record
Selected · Sources reviewed 24 Aug 2026
Stateful Earth-observation work across executable tools, temporal retrieval and cross-modal evidence.
Checks: Pass@k, Tool-any, illegal-call rate, and task-category scoring
Read the benchmark record