The public site began with a dated October 1 snapshot and a static portal roster. A direct inspection of the machines, running launch arguments, active gateway configs and read-only dashboard endpoints showed that parts of that description were already stale.
Hardware verified directly
- NVIDIA DGX Spark and ASUS GX10, both GB10 unified-memory systems. The old site described the GX10 as a second DGX Spark.
- Mac Studio M5 Max, 18 CPU cores, 40 GPU cores, 48 GB unified memory.
- RTX 5090 workstation, Ryzen 9 9950X3D, Windows + WSL2. The GPU driver reports 32,607 MiB device capacity. Guest RAM is not presented as installed Windows RAM.
Some dashboard hardware labels also lag the real machines. System reports take precedence for hardware facts; active configs take precedence for role placement.
The current placement
The primary coder is Qwen3.8-Flash-Next NVFP4 on vLLM 0.30.0, across both GB10 nodes with tensor and expert parallelism. Its running launch arguments specify 262,144 context, eight sequences, an 8,192 prefill batch, allocator 0.74 and three native MTP draft positions.
The coordinator now routes agent and embeddings to the Studio. The Studio runs an exclusive primary swap group for its large models and image alternatives, with embeddings in a separate persistent group.
The workstation is the primary route for media and several utility roles: OCR, translation, transcription and speech. Several coding candidates also exist in its configured inventory. Configured does not mean benchmark-qualified or loaded.
Keep the evidence separate
No new inference benchmark was run during this read-only audit. These are configuration facts, not throughput claims. The older Spark studies and September Studio evaluation keep their own dates and protocols. Upstream deployment-kit changelog measurements were not republished as results from this lab.
The reusable lessons and current limits are collected in hardware, setup guides and the software stack.