Skip to content

Stack

What actually runs

Layers that actually run in the lab, dated. No private URLs, ports, or launch flags. Rows drop off when they stop being true.

Stack layers motif STACK LAYERS apps → api → swap → silicon Apps OpenAI-compatible API llama-swap Models / Hardware L1L2L3L4
Runtime path. Inventory below is dated from the lab notebook.
  1. Hardware

    • DGX Spark / GB10

      Two desktop Sparks; NVIDIA GB10 for local TP2 inference.

      as of 2026-10-01

    • Mac Studio / M5 Max

      Full-stack local-AI box in the same pool. Embeddings at the Oct 1 snapshot.

      as of 2026-10-01

    • Gaming PC / RTX 5090

      Windows host; GPU exposed through WSL. On-demand OCR / translate / speech.

      as of 2026-10-01

  2. Drivers

    • Ubuntu NVIDIA kernel

      Spark OS kernel for the GB10 boxes.

      as of 2026-09-26

    • NVIDIA driver 580 / CUDA 13

      Driver and CUDA line seen on the Sparks.

      as of 2026-09-26

  3. Inference

    • vLLM

      Engine under dual-Spark TP2 coder-flash at the 2026-10-01 snapshot.

      as of 2026-10-01

    • SGLang

      Inference server used on the Sparks as of the prior dated row.

      as of 2026-09-26

    • Qwen3.8-Flash-Next NVFP4

      coder / coder-flash / vision · TP2 both Sparks · 262144 context.

      as of 2026-10-01

    • Qwen 3.8 27B NVFP4 + DFlash

      Prior coder model with a speculative DFlash draft.

      as of 2026-09-26

  4. Gateway

    • llama-swap

      Private trusted-network gateway; OpenAI + Anthropic Messages compatible; pools Sparks, Studio, and the 5090.

      as of 2026-10-01

  5. Side services

    • Open WebUI

      Local chat UI. Not public.

      as of 2026-09-26

    • Kokoro

      Local voice synthesis. Not public.

      as of 2026-09-26

    • SearxNG

      Local search. Not public.

      as of 2026-09-26

    • Prometheus / Grafana

      Lab metrics. Not a public status page. No live charts published here.

      as of 2026-10-01