Research protocol

How the Arena reads evidence

A result is a reproducible chain of inputs, complete configuration, execution, artifacts, verifier checks, recovery behaviour, and candid failure states.

Complete comparison unit

A result belongs to a complete configuration: model, provider, prompt, harness, tools, runtime, retry policy, and pinned benchmark, dataset, and verifier versions. A model name alone is not comparable.

Deterministic verification

Verifiers inspect output content and independently recompute expected values where possible. Empty files, fabricated numbers, unsafe spatial operations, and whole-grid shortcuts are rejected. An LLM judge is never the sole correctness signal.

Inspectable artifacts

Published records retain configuration identity, attempts, generated code, execution logs, artifact links where available, and verifier detail. Reproducibility is a reported condition, not an assumption.

Measures reported together

Completion is read alongside cost, latency, retry or recovery behaviour, false-success rejections, and reproducibility. A fast plausible answer is not treated as accepted work.

False-success classes

a plausible map can still be wrong

Artifact and area failures

Empty or zero-pixel output, area-of-interest drift, unsafe buffering, and raster misalignment can look complete while failing the ground truth.

Spatial meaning failures

CRS or unit error, nodata or cloud-mask failure, and a plausible but spatially wrong map are correctness failures, not cosmetic defects.

Evidence failures

A provenance break or an unrepeatable result cannot support acceptance, even when an output looks convincing.

Experimental boundary

Historical Cloudflare rows are unreconciled. Only current, inspectable rows may appear as an experimental snapshot; absent rows are not compared or inferred.