Skip to content

Fleet

Four boxes. One pool.

The four public machines pool through llama-swap — a private trusted-network AI gateway. Model swap, routing, and one OpenAI- plus Anthropic-compatible front door on the lab side. No public URL. No private hostnames.

as of 2026-10-01 13:51:36 CDT

Lab snapshot

Headline stats

2026-10-01 13:51:36 CDT

As of 2026-10-01 13:51:36 CDT, dual-Spark Qwen3.8-Flash-Next NVFP4 (coder / coder-flash / vision) was loaded at 262144 context with about 63 GiB GPU memory per Spark worker. Live generation was idle (0 tok/s).

Attributed lab snapshot from the private trusted-network gateway. Not a live dashboard and not a public API.

  • Loaded roles

    coder-flash · coder · vision

    Qwen3.8-Flash-Next NVFP4 · TP2 both Sparks · vLLM · 262144 context

  • Context

    262144

    coder-flash context window at snapshot

  • GPU memory

    ~63 GiB / worker

    ~50% of ~124.6 GiB unified pool on each Spark

  • GPU util

    ~92–96%

    Sparks during TP2 residency. Snapshot gauge, not a live chart.

  • Live generation

    idle · 0 tok/s

    generationTps was 0 on every node in the probe window.

  • On-demand

    ocr · translate · stt · tts

    Selectors available; unloaded at snapshot. Gaming PC peer online.

Live decode throughput was not observed. Catalog and study tok/s figures elsewhere on this site are dated page copy, not this snapshot.

Fleet topology — llama-swap pooling hub FLEET TOPOLOGY llama-swap pooling llama-swap Spark A Spark B Mac Studio Gaming PC
Four public machines. Private trusted-network llama-swap gateway. No private hostnames.

dgx-spark-a

DGX Spark A

Tensor-parallel inference (TP2)

TP2 resident · idle

GPU
NVIDIA GB10
Memory
Unified GB10 pool. Snapshot 2026-10-01 13:51:36 CDT: ~63 GiB GPU resident (~50% of ~124.6 GiB unified).
Storage
Local NVMe. Capacity class only — live fill is ops data.
Notes
One of two DGX Spark boxes. TP2 worker for coder-flash / coder / vision (Qwen3.8-Flash-Next NVFP4, 262144 context). GPU util ~92% during that residency. Live generation idle (0 tok/s).

as of 2026-10-01 13:51:36 CDT

dgx-spark-b

DGX Spark B

Tensor-parallel inference (TP2)

TP2 resident · idle

GPU
NVIDIA GB10
Memory
Same unified GB10 class as Spark A. Snapshot: ~63 GiB GPU resident (~50% of ~124.6 GiB unified).
Storage
Local NVMe. Capacity class only — live fill is ops data.
Notes
Second Spark. TP2 peer for the same dual-Spark coder-flash residency. GPU util ~96% at snapshot. Live generation idle (0 tok/s).

as of 2026-10-01 13:51:36 CDT

mac-studio

Mac Studio

Full-stack / embeddings

embed ready · agent unloaded

GPU
Apple M5 Max
Memory
Unified Apple memory. Live GPU gauges from the Oct 1 pull were mis-templated and are not published.
Storage
Local SSD. Size not published.
Notes
Full-stack local-AI machine in the llama-swap pool. studio/embed (Qwen3-Embedding-4B) was ready at the Oct 1 snapshot. agent / lightning selectors exist; they were unloaded live. Catalog ~80 tok/s figures are dated page copy, not live telemetry.

as of 2026-10-01 13:51:36 CDT

gaming-pc

Gaming PC

On-demand vision / speech

online · on-demand unloaded

GPU
NVIDIA RTX 5090 via WSL on Windows
Memory
Discrete RTX 5090. Snapshot: peer online, ~3% VRAM, no local model loaded.
Storage
Local disk. Size not published.
Notes
Windows host, inference through WSL, pooled with the Sparks and Studio via llama-swap. OCR, translate, STT, and TTS selectors pin here on demand. All unloaded at the 2026-10-01 13:51:36 CDT snapshot.

as of 2026-10-01 13:51:36 CDT