Journal
2026-07-19
Multimodal creative agent → Gemma-4-26B QAT
Lab study note from the internal benches record (page last-modified 2026-09-12). Not live Oct 1 telemetry.
Which model should be the always-on agent? Four candidates ran the same long-form, tool, vision, cache, concurrency, and 120K-context suite. The current Agent remained the strongest operational baseline.
Verdict — keep Agent authoritative
Gemma-4-26B QAT won the job: best blind writing average, MoE-class latency, full context retrieval, no reliability regression.
- Blind overall: 8.35 / 10 (six constrained assignments)
- Tools: 91.2%
- Single stream: 60.4 tok/s
- Four-way aggregate: 129.9 tok/s
- Warm 17K prefix: 0.303 s (23.2× over cold)
- Deep edge pair: 78.3 s · 119,593 prompt tokens
Rejected
- Goetia (Gemma-4 heretic): tools and vision matched; blind writing 6.65. Operationally strong, not the writer.
- Qwen3.6 Fable Fusion: tools 97.1%, but blind 4.77 and ~10.8 tok/s. Rejected for the always-on seat.
- Ortenzya (dense 31B): blind 8.17, interesting as a routed creative specialist at 9.8 tok/s — not the resident agent.
Production restored with seven aliases live. This is dated study-page content, not the Oct 1 idle snapshot.