Skip to content
← All setup guides

Guide 06 / Measurement

Measure the experience, not just tok/s

A repeatable protocol for latency, decode, context, concurrency, correctness, and honest comparisons.

2 MIN READ / REVIEWED 02 OCT 2026

01Freeze the configuration and workload

Record hardware, artifact revision, quantization, engine, chat template, thinking mode, sampling, prompt length, output cap and concurrency. Separate cold startup from a loaded model and a cached prefix. Run one contender at a time unless contention is the experiment.

02Report distinct measurements

Missing measurements stay missing. A zero idle generation gauge is not a model performance result. Catalog estimates and configured limits do not become measured results.

  • TTFT: request start to first output; include prefill and queueing context.
  • Decode: output positions per second after generation begins. State the timing source and denominator.
  • Wall time: time to finish the real task, including loading, tools and validation.
  • Concurrent throughput: aggregate output rate plus per-stream latency and failures.
  • Capacity: memory, cancellations, admission behavior and retrieval near the tested context boundary.

03Use repeatable correctness checks

Keep raw failures, task counts and scoring rules. Repeat noisy custom gates and report medians and spread. Compile generated code or run executable tests where appropriate; regex framework checks alone cannot establish correctness. HumanEval chat adaptation, local tool suites and public leaderboard protocols must be identified separately.

04Compare within a shared protocol

Our Studio chart compares three models on the same September machine and short-probe protocol. It is not a Spark-versus-Mac hardware ranking. Model, runtime, quantization, prompts and dates differ across those studies. The older Spark tables remain separate.

A minimal result record

{
  "date": "YYYY-MM-DD",
  "hardware": "device, memory, driver",
  "artifact": "model revision and quantization",
  "runtime": "version or image digest",
  "protocol": "sampling, thinking, input/output lengths",
  "concurrency": 1,
  "samples": [],
  "correctness": "suite and denominator",
  "limitations": "what this result does not establish"
}

05Make a promotion reversible

A challenger needs a useful quality, latency, reliability or capacity improvement for its assigned role. Preserve the incumbent, rerun after runtime changes and keep a rollback path. The July Laguna reversal is the reason this journal records rejected and superseded decisions alongside wins.

Keep going

Related guides