Setup guides / from first model to a working system
Build a baseline you can trust.
Practical paths through hardware, runtimes, memory, routing, and measurement. Based on the lab’s configuration and the mistakes its experiments exposed.
New to local inference?
Start with one useful workload.
Choose an artifact, budget memory, and qualify one backend before building a fleet.
7 OF 7 GUIDES
Your first useful local LLM
Start with one workload, one model, and a memory budget. Expand after a measured baseline.
2 MIN READ / UPDATED OCT 2026
DGX Spark & GX10: capacity with discipline
ARM64, unified memory, distributed inference, and the actual settings on our two-GB10 coder.
2 MIN READ / UPDATED OCT 2026
RTX 5090 on Windows + WSL2
Verify the Windows driver from Linux, enable CUDA, and budget a 32 GB GPU for useful work.
2 MIN READ / UPDATED OCT 2026
Mac Studio: MLX, Metal & memory ownership
Run local models on Apple silicon without confusing installed memory, configured context, and verified capacity.
2 MIN READ / UPDATED OCT 2026
Routing is easy. Model lifecycle is the work.
Persistence, swap groups, alias drift, child processes, and protocol fidelity across heterogeneous machines.
2 MIN READ / UPDATED OCT 2026
Measure the experience, not just tok/s
A repeatable protocol for latency, decode, context, concurrency, correctness, and honest comparisons.
2 MIN READ / UPDATED OCT 2026
The problems behind a model that “works”
Lost tool calls, memory that survives unloading, untested context, bad gauges, and regressions.
2 MIN READ / UPDATED OCT 2026