Skip to content

Journal

2026-08-10

Muse Glimmer dual-role reject

Lab study note from the internal benches record (page last-modified 2026-09-12). Not live Oct 1 telemetry.

Day-0 open weights from Meta Superintelligence Labs (Apache 2.0). Dense multimodal 30B distilled from Muse Spark. Question: is it a fleet fit as coder or always-on agent?

Verdict — rejected both roles

  • Not a coding lane. devcode 81.5% vs champion 86.4% (−4.9, outside ±4 noise). devtool collapsed to 25.0%. Wall-clock 20× slower. Native context 131K fails the verified-256K coder gate.
  • Not always-on agent. Tools 91.2% tie the Gemma agent, and developer sidecar 30/30 matches Omni. Dense 30B at 10.1 tok/s loses to MoE Gemma at 60.4 tok/s on every felt agent loop.
  • Not justified as an on-demand specialist. Ornith already owns tools, Omni owns developer vision, Gemma owns fast agent.

Coder axes vs Coder-Next

  • c1: 10.1 vs 50.9 tok/s (0.20×). Even DFlash code peak 34.3 tok/s stays 0.67×.
  • GPU footprint ~21 GiB vs 53.9 GiB — a win that does not pay for 0.2× decode.
  • Cold start 17 s vs 90 s — also not enough.

Same promotion rule as the Laguna revert: do not promote on tool quality alone when single-stream is 0.2×. Not added to llama-swap.

← All notes