Skip to content
← All setup guides

Guide 04 / Mac Studio

Mac Studio: MLX, Metal & memory ownership

Run local models on Apple silicon without confusing installed memory, configured context, and verified capacity.

2 MIN READ / REVIEWED 02 OCT 2026

01Start from the measured device

This Studio has an M5 Max, 18 CPU cores, 40 GPU cores and 48 GB installed unified memory. The September evaluation recorded a recommended Metal working set of 37.44 GiB. Leave room for caches, runtime buffers and the OS rather than assigning the installed total to model weights.

02Use the artifact your engine expects

MLX-LM runs supported Apple-silicon models and exposes generation, chat, cache and prefill controls. Our evaluated Lightning and Qwen artifacts use MLX 4-bit; Gemma and embeddings use GGUF through llama.cpp Metal. These are different inference paths.

MLX reference quickstart

# Example isolated environment using the evaluated MLX versions
python3 -m venv .venv
source .venv/bin/activate
python -m pip install mlx==0.32.2 mlx-lm==0.31.3
mlx_lm.generate --help

# Set this to a model artifact you have reviewed
LOCAL_MODEL_ID="your-reviewed-mlx-model"
mlx_lm.generate --model "$LOCAL_MODEL_ID" --prompt "Explain KV cache briefly."

03Own the backend and the proxy together

The original Studio launcher could stop its proxy and leave the MLX backend alive. The repaired supervisor owns both children and reclaims both on unload. Test this using process and memory observations after an actual swap; a successful unload response is insufficient.

04Keep embeddings independent

The current primary group swaps Lightning, Qwen, both Gemma roles and image alternatives exclusively. Embeddings use a separate persistent group. TTL zero alone did not establish persistence in the original configuration.

05Read the evaluation at its measured scale

Qwen passed 156 of 164 chat-adapted HumanEval tasks; Lightning passed 140. Gemma was not run on HumanEval. Short loaded-backend code probes had median decode of 33.0, 150.1 and 78.1 tok/s respectively. The synthetic retrieval probes reached about 34K input positions; they did not qualify 64K or the advertised 262K ceiling.

Primary references

Upstream docs describe supported behavior. Lab-specific configurations and numeric results retain their inspection or experiment dates.

Keep going

Related guides