Complete comparison unit
A result belongs to a complete configuration: model, provider, prompt, harness, tools, runtime, retry policy, and pinned benchmark, dataset, and verifier versions. A model name alone is not comparable.
Research protocol
A result is a reproducible chain of inputs, complete configuration, execution, artifacts, verifier checks, recovery behaviour, and candid failure states.
A result belongs to a complete configuration: model, provider, prompt, harness, tools, runtime, retry policy, and pinned benchmark, dataset, and verifier versions. A model name alone is not comparable.
Verifiers inspect output content and independently recompute expected values where possible. Empty files, fabricated numbers, unsafe spatial operations, and whole-grid shortcuts are rejected. An LLM judge is never the sole correctness signal.
Published records retain configuration identity, attempts, generated code, execution logs, artifact links where available, and verifier detail. Reproducibility is a reported condition, not an assumption.
Completion is read alongside cost, latency, retry or recovery behaviour, false-success rejections, and reproducibility. A fast plausible answer is not treated as accepted work.
Empty or zero-pixel output, area-of-interest drift, unsafe buffering, and raster misalignment can look complete while failing the ground truth.
CRS or unit error, nodata or cloud-mask failure, and a plausible but spatially wrong map are correctness failures, not cosmetic defects.
A provenance break or an unrepeatable result cannot support acceptance, even when an output looks convincing.
Historical Cloudflare rows are unreconciled. Only current, inspectable rows may appear as an experimental snapshot; absent rows are not compared or inferred.