Journal
2026-07-23
Laguna SM121 runtime
Lab study note from the internal benches record (page last-modified 2026-09-12). Not live Oct 1 telemetry.
Question: can Laguna’s serving runtime be tuned on GB10? Sweep covered DFlash, FP8 KV, CUDA graphs, and the allocator.
Outcome
- Promoted (this day): eager mode with a .79 allocator. .80 crossed the earlyoom threshold during startup.
- Rejected: DFlash K=7. Draft accepted 1,320 of 8,666 tokens — 15.2% — and reduced c4 throughput.
- Best profile still only 17.6 tok/s c1 / 54.6 c4. KV pool allocated 1,173,283 tokens (4.48× the configured 262,144 window).
The ceiling was the model on this hardware, not the configuration. This is the promotion the July 26 same-node A/B reversed.