This is a historical local measurement from the Spark tuning review, inspected again on October 2. Not live Oct 1 telemetry. External reference-rig numbers in that review are excluded from this public result.
Same model, different acceptance
Qwen3.6-35B-A3B Q6_K on llama.cpp b9967, one slot, quiet hardware. Native MTP used six draft positions.
- Code, temperature 0.3: 64.1 → 100.9 tok/s with MTP enabled.
- Creative writing, temperature 0.7: 64.6 → 53.6 tok/s with MTP enabled.
The local summary attributes the difference to draft acceptance: rejected speculative work can cost more than it saves. Repeat count and spread were not recorded in the reviewed summary, so this is not a universal speedup claim.
Configure by role
Keep speculation as a measured per-role choice. A result on this model and sampling policy does not transfer automatically to a different checkpoint, drafting method or runtime. The current Flash-Next MTP3 configuration is a separate record.