Journal
2026-07-21
Three-lane coder and the July 21 gates
Lab study note from the internal benches record (page last-modified 2026-09-12). Not live Oct 1 telemetry.
Five controlled studies landed the same day.
Three-lane coder — deployed
Capacity-aware dispatcher. Two Spark slots plus one local lane. Four concurrent 256-token streams reached 99.6 aggregate tok/s (42.0–56.8 per stream); excess queued, no 409s. Session affinity keeps a conversation on one worker so its prompt cache stays local.
Coder 2 × 256K — qualified
Will two independent 262,144-token slots hold on one Spark? Yes. Both deep sessions retrieved correctly near 241K. 52.6 tok/s single · 76.0 aggregate on two streams · warm turns 4.3–4.4 s. About 55.2 GiB. This is warm-session affinity, not a 2× cold accelerator.
Laguna S vs coder — rejected
First evaluation, including native TP=2. Code 79.7% vs incumbent 87.8%. Native TP=2 deadlocked on 512-token generations. Not a quality or throughput upgrade.
Coder quantization — keep UD-Q4_K_XL
Three Q5 variants beside production Unsloth dynamic Q4, same 256K settings.
- UD-Q4_K_XL (keep): 89.2% code, 50.1 tok/s, 49.6 GB
- Q5_K_M: tools rose to 66.7%, code fell to 86.2%
- UD-Q5_K_M: 87.4% code, ~17% slower
- UD-Q5_K_XL: best Swift 74.1%, still lost overall quality and speed
More bits did not make a better default. Q5 cost 7–10 GB and ~17% decode for no complete-workload win.
Split vision router — Nemotron Omni
Images are described once and replaced with cached text; coding stays on the text-only coder.
- Keep vision: Nemotron Omni — 30/30 developer facts in 10.0 s; production context raised to 262,144. Perfect retrieval through 240,123 actual tokens.
- Keep coder: Qwen3-Coder-Next — 87.6% code, 75.0% developer tools, retrieval near 255K.
- Fallback: Qwen3.5-9B — also 30/30, but 80.7 s wall (8×).
- Does not fit: Nemotron Super 120B — loaded, then OOM-killed.