01Define the job before the model
Choose a concrete first task: code review, interactive chat, document embeddings, or media generation. Select a model supported by your runtime, with a license that permits your use. Download from the model publisher or a reviewed conversion, and record the artifact revision.
- Apple silicon: MLX-format checkpoints or GGUF with Metal.
- NVIDIA: a CUDA-compatible runtime and model architecture.
- GGUF, MLX 4-bit, NVFP4 and FP8 are different deployment artifacts; they are not interchangeable.
02Make a conservative memory budget
Reserve memory for the OS, model weights, KV cache, runtime buffers and concurrent requests. Active MoE parameters describe compute, not the complete checkpoint footprint. A model that loads once has not passed an operating-capacity test.
03Bring up one backend first
Use the relevant hardware guide below. Confirm chat formatting, a tool call if needed, cancellation, and a warm second turn directly against the backend. Add a gateway after that baseline; check that its answers and response fields survive the extra hop.
04Keep a small deployment manifest
Pin your runtime and model revision, preserve a known-good rollback, and record the workload you used to qualify the configuration.
Example manifest · fill in your own values
hardware: your measured device and memory
model: publisher + artifact revision
runtime: version or image digest
quantization: actual artifact format
context_budget: chosen initial limit
active_requests: 1
checks: [chat, tools, cancellation, warm_turn, memory]
rollback: last qualified configurationPrimary references
Upstream docs describe supported behavior. Lab-specific configurations and numeric results retain their inspection or experiment dates.