01Freeze the configuration and workload
Record hardware, artifact revision, quantization, engine, chat template, thinking mode, sampling, prompt length, output cap and concurrency. Separate cold startup from a loaded model and a cached prefix. Run one contender at a time unless contention is the experiment.
02Report distinct measurements
Missing measurements stay missing. A zero idle generation gauge is not a model performance result. Catalog estimates and configured limits do not become measured results.
- TTFT: request start to first output; include prefill and queueing context.
- Decode: output positions per second after generation begins. State the timing source and denominator.
- Wall time: time to finish the real task, including loading, tools and validation.
- Concurrent throughput: aggregate output rate plus per-stream latency and failures.
- Capacity: memory, cancellations, admission behavior and retrieval near the tested context boundary.
03Use repeatable correctness checks
Keep raw failures, task counts and scoring rules. Repeat noisy custom gates and report medians and spread. Compile generated code or run executable tests where appropriate; regex framework checks alone cannot establish correctness. HumanEval chat adaptation, local tool suites and public leaderboard protocols must be identified separately.
05Make a promotion reversible
A challenger needs a useful quality, latency, reliability or capacity improvement for its assigned role. Preserve the incumbent, rerun after runtime changes and keep a rollback path. The July Laguna reversal is the reason this journal records rejected and superseded decisions alongside wins.