Full setup
A result records the model, prompt, tools, retry policy, task set, data, and verifier version. Model names alone do not make results comparable.
Research protocol
Arena records the task, model, output, checks, time, cost, and failures. A result is a reproducible chain of inputs, complete configuration, execution, artefacts, verifier checks, recovery behaviour, and candid failure states.
A result records the model, prompt, tools, retry policy, task set, data, and verifier version. Model names alone do not make results comparable.
Verifiers inspect the output and recompute expected values where possible. Empty files, fabricated numbers, unsafe spatial operations, and shortcuts fail.
Each record keeps the output, attempts, logs, artefacts where available, and verifier detail. Reproducibility is a reported condition, not an assumption.
A result also shows cost, latency, retries, false-success rejections, and reproducibility. A fast plausible answer is not accepted work.
The verifier's deterministic checks all passed for the final attempt within the task budget.
The output exists but failed one or more independent checks. The failing check names expected and actual values.
No executable attempt was produced: an inaccessible model, a sandbox start failure, or an empty program.
A run mixed accepted and non-accepted trials. Headline scores quote the accepted subset and list every run beside it.
A run stopped before completing its planned trials, or exceeded its wall-clock budget. Partial rows remain visible.
Empty output, area drift, unsafe buffers, and raster misalignment can look complete while failing the ground truth.
Wrong CRS, units, nodata handling, or cloud masks are correctness failures, not style differences.
Invented statistics, hardcoded expected values, or outputs written without reading inputs. Independent recomputation catches these.
An unrepeatable result cannot support acceptance, even when the output looks convincing.
Run ea17d555 shows accepted lanes beside rejected ones. Run c55f0f7f shows a mixed outcome kept visible as inconclusive rather than averaged away.
Every configuration runs at least three trials per task before any reliability claim. Single observations are labelled as such. Pass-at-k is reported when retries are allowed.
A judgement model can rate plausibility but cannot prove geometric correctness. Its scores never contribute to acceptance and appear, if at all, as descriptive context next to deterministic results.
When a run finishes, its comparison table is written once as a content-addressed snapshot. Published figures read from sealed snapshots so they cannot drift afterwards.
Arena publishes only current records that can be inspected. Incomplete records are not used for comparisons.