01Start from the measured device
This Studio has an M5 Max, 18 CPU cores, 40 GPU cores and 48 GB installed unified memory. The September evaluation recorded a recommended Metal working set of 37.44 GiB. Leave room for caches, runtime buffers and the OS rather than assigning the installed total to model weights.
02Use the artifact your engine expects
MLX-LM runs supported Apple-silicon models and exposes generation, chat, cache and prefill controls. Our evaluated Lightning and Qwen artifacts use MLX 4-bit; Gemma and embeddings use GGUF through llama.cpp Metal. These are different inference paths.
MLX reference quickstart
# Example isolated environment using the evaluated MLX versions
python3 -m venv .venv
source .venv/bin/activate
python -m pip install mlx==0.32.2 mlx-lm==0.31.3
mlx_lm.generate --help
# Set this to a model artifact you have reviewed
LOCAL_MODEL_ID="your-reviewed-mlx-model"
mlx_lm.generate --model "$LOCAL_MODEL_ID" --prompt "Explain KV cache briefly."03Own the backend and the proxy together
The original Studio launcher could stop its proxy and leave the MLX backend alive. The repaired supervisor owns both children and reclaims both on unload. Test this using process and memory observations after an actual swap; a successful unload response is insufficient.
04Keep embeddings independent
The current primary group swaps Lightning, Qwen, both Gemma roles and image alternatives exclusively. Embeddings use a separate persistent group. TTL zero alone did not establish persistence in the original configuration.
05Read the evaluation at its measured scale
Qwen passed 156 of 164 chat-adapted HumanEval tasks; Lightning passed 140. Gemma was not run on HumanEval. Short loaded-backend code probes had median decode of 33.0, 150.1 and 78.1 tok/s respectively. The synthetic retrieval probes reached about 34K input positions; they did not qualify 64K or the advertised 262K ceiling.
Primary references
Upstream docs describe supported behavior. Lab-specific configurations and numeric results retain their inspection or experiment dates.