Stack
What actually runs
Layers that actually run in the lab, dated. No private URLs, ports, or launch flags. Rows drop off when they stop being true.
Hardware
DGX Spark / GB10
Two desktop Sparks; NVIDIA GB10 for local TP2 inference.
as of 2026-10-01
Mac Studio / M5 Max
Full-stack local-AI box in the same pool. Embeddings at the Oct 1 snapshot.
as of 2026-10-01
Gaming PC / RTX 5090
Windows host; GPU exposed through WSL. On-demand OCR / translate / speech.
as of 2026-10-01
Drivers
Ubuntu NVIDIA kernel
Spark OS kernel for the GB10 boxes.
as of 2026-09-26
NVIDIA driver 580 / CUDA 13
Driver and CUDA line seen on the Sparks.
as of 2026-09-26
Inference
vLLM
Engine under dual-Spark TP2 coder-flash at the 2026-10-01 snapshot.
as of 2026-10-01
SGLang
Inference server used on the Sparks as of the prior dated row.
as of 2026-09-26
Qwen3.8-Flash-Next NVFP4
coder / coder-flash / vision · TP2 both Sparks · 262144 context.
as of 2026-10-01
Qwen 3.8 27B NVFP4 + DFlash
Prior coder model with a speculative DFlash draft.
as of 2026-09-26
Gateway
llama-swap
Private trusted-network gateway; OpenAI + Anthropic Messages compatible; pools Sparks, Studio, and the 5090.
as of 2026-10-01
Side services
Open WebUI
Local chat UI. Not public.
as of 2026-09-26
Kokoro
Local voice synthesis. Not public.
as of 2026-09-26
SearxNG
Local search. Not public.
as of 2026-09-26
Prometheus / Grafana
Lab metrics. Not a public status page. No live charts published here.
as of 2026-10-01