Skip to content

Journal

2026-10-01

Snapshot, 1 October 2026

Attributed 2026-10-01 13:51:36 CDT. Dual Sparks had coder-flash / coder / vision loaded — Qwen3.8-Flash-Next NVFP4, 262144 context, TP2 across both GB10 boxes. About 63 GiB GPU memory per Spark worker (~50% of a ~124.6 GiB unified pool). GPU util sat around 92–96% from that residency.

Live generation was idle (0 tok/s) on every node. That is the only live throughput number from this pull. Catalog figures such as ~80 tok/s for agent / lightning are dated page copy from the Sep 12 2026 benches page; those selectors were unloaded.

OCR, translate, STT, and TTS were available as on-demand selectors and unloaded. The Gaming PC (RTX 5090) peer was online. Mac Studio M5 Max had embeddings ready; its GPU block from that pull was mis-templated and is not published.

No private hostnames, no live charts, no public gateway URL.

← All notes