Skip to content
← All setup guides

Guide 01 / Fundamentals

Your first useful local LLM

Start with one workload, one model, and a memory budget. Expand after a measured baseline.

2 MIN READ / REVIEWED 02 OCT 2026

01Define the job before the model

Choose a concrete first task: code review, interactive chat, document embeddings, or media generation. Select a model supported by your runtime, with a license that permits your use. Download from the model publisher or a reviewed conversion, and record the artifact revision.

  • Apple silicon: MLX-format checkpoints or GGUF with Metal.
  • NVIDIA: a CUDA-compatible runtime and model architecture.
  • GGUF, MLX 4-bit, NVFP4 and FP8 are different deployment artifacts; they are not interchangeable.

02Make a conservative memory budget

Reserve memory for the OS, model weights, KV cache, runtime buffers and concurrent requests. Active MoE parameters describe compute, not the complete checkpoint footprint. A model that loads once has not passed an operating-capacity test.

03Bring up one backend first

Use the relevant hardware guide below. Confirm chat formatting, a tool call if needed, cancellation, and a warm second turn directly against the backend. Add a gateway after that baseline; check that its answers and response fields survive the extra hop.

04Keep a small deployment manifest

Pin your runtime and model revision, preserve a known-good rollback, and record the workload you used to qualify the configuration.

Example manifest · fill in your own values

hardware: your measured device and memory
model: publisher + artifact revision
runtime: version or image digest
quantization: actual artifact format
context_budget: chosen initial limit
active_requests: 1
checks: [chat, tools, cancellation, warm_turn, memory]
rollback: last qualified configuration

Primary references

Upstream docs describe supported behavior. Lab-specific configurations and numeric results retain their inspection or experiment dates.

Keep going

Related guides