The September 22–23 evaluation ran sequential large-model tests on an M5 Max Mac Studio: 18 CPU cores, 40 GPU cores and 48 GB unified memory. This entry is a curated summary of the reviewed scorecard and numeric records, not a new October benchmark.
Three models, three jobs
Lightning MLX 4-bit was the fast-agent candidate. Qwen3.8-27B MLX 4-bit was the coder. Gemma-4 QAT Q4_K_M was evaluated as a writer through llama.cpp Metal.
- Chat-adapted HumanEval: Qwen 156/164 (95.1%), Lightning 140/164 (85.4%). Gemma was not run on this coding benchmark.
- Code decode medians: Lightning 150.1 tok/s, Qwen 33.0 tok/s, Gemma 78.1 tok/s. Each is a median of three short 512-output probes on a loaded backend.
- Framework criteria: 71.5% / 89.5% / 84.3%. These are regex checks, not compilation results.
- Prose constraints: 73.5% / 87.2% / 97.0%. Constraint adherence is not a literary-quality rating.
The serving path changed the answer
The original metrics proxy reconstructed content and dropped non-streaming tool calls. Separate stochastic tool runs scored 32.4% through that proxy, 79.4% directly and 88.2% through the repaired path. The broken response structure was demonstrated; the difference between the latter two runs is not proof of a model quality gain.
The original launcher also left the MLX backend outside normal unload ownership. The repaired supervisor owns backend and proxy together. Embeddings gained an independent persistent group; TTL zero by itself had not kept them resident. Passive metadata polling was changed so it did not wake inactive models.
What the measurements establish
HumanEval used all 164 original tasks, one greedy chat sample, a 2048-output cap, thinking off, sandboxed executable tests and no answer repair. All reference solutions passed runner validation. This is chat-adapted pass@1, not EvalPlus, SWE-bench or a vendor leaderboard reproduction. Training-data contamination remains possible.
The local tools and prose suites are custom gates. Some developer-tool regex checks penalized valid alternatives, so their raw percentages are not autonomous coding success rates. The retrieved synthetic records reached about 34K input positions, not a 64K or 262K capacity qualification. Manual Swift review found correctness issues even in answers that scored well.
The public benchmark records retain these distinctions.