Skip to content

Journal

2026-07-19

Multimodal creative agent → Gemma-4-26B QAT

Lab study note from the internal benches record (page last-modified 2026-09-12). Not live Oct 1 telemetry.

Which model should be the always-on agent? Four candidates ran the same long-form, tool, vision, cache, concurrency, and 120K-context suite. The current Agent remained the strongest operational baseline.

Verdict — keep Agent authoritative

Gemma-4-26B QAT won the job: best blind writing average, MoE-class latency, full context retrieval, no reliability regression.

  • Blind overall: 8.35 / 10 (six constrained assignments)
  • Tools: 91.2%
  • Single stream: 60.4 tok/s
  • Four-way aggregate: 129.9 tok/s
  • Warm 17K prefix: 0.303 s (23.2× over cold)
  • Deep edge pair: 78.3 s · 119,593 prompt tokens

Rejected

  • Goetia (Gemma-4 heretic): tools and vision matched; blind writing 6.65. Operationally strong, not the writer.
  • Qwen3.6 Fable Fusion: tools 97.1%, but blind 4.77 and ~10.8 tok/s. Rejected for the always-on seat.
  • Ortenzya (dense 31B): blind 8.17, interesting as a routed creative specialist at 9.8 tok/s — not the resident agent.

Production restored with seven aliases live. This is dated study-page content, not the Oct 1 idle snapshot.

← All notes