Journal
2026-08-10
Muse Glimmer dual-role reject
Lab study note from the internal benches record (page last-modified 2026-09-12). Not live Oct 1 telemetry.
Day-0 open weights from Meta Superintelligence Labs (Apache 2.0). Dense multimodal 30B distilled from Muse Spark. Question: is it a fleet fit as coder or always-on agent?
Verdict — rejected both roles
- Not a coding lane. devcode 81.5% vs champion 86.4% (−4.9, outside ±4 noise). devtool collapsed to 25.0%. Wall-clock 20× slower. Native context 131K fails the verified-256K coder gate.
- Not always-on agent. Tools 91.2% tie the Gemma agent, and developer sidecar 30/30 matches Omni. Dense 30B at 10.1 tok/s loses to MoE Gemma at 60.4 tok/s on every felt agent loop.
- Not justified as an on-demand specialist. Ornith already owns tools, Omni owns developer vision, Gemma owns fast agent.
Coder axes vs Coder-Next
- c1: 10.1 vs 50.9 tok/s (0.20×). Even DFlash code peak 34.3 tok/s stays 0.67×.
- GPU footprint ~21 GiB vs 53.9 GiB — a win that does not pay for 0.2× decode.
- Cold start 17 s vs 90 s — also not enough.
Same promotion rule as the Laguna revert: do not promote on tool quality alone when single-stream is 0.2×. Not added to llama-swap.