Skip to content

Journal

2026-07-26

Coder revert to Coder-Next

Lab study note from the internal benches record (page last-modified 2026-09-12). Not live Oct 1 telemetry.

The first same-node, same-day A/B of the two coders, one resident at a time. The July 23 promotion of Laguna had been made on a single axis — code score — and every other axis favoured the incumbent. On this day the code score reversed too.

Measured (Laguna S 2.1 vs Qwen3-Coder-Next)

  • devcode: 87.2% vs 90.9% (Coder-Next +3.7)
  • Same-suite wall: 435.7 s vs 82 s (5.3×)
  • devtool: 75.0% vs 66.7% — Laguna’s one durable edge
  • c1: 17.7 vs 52.7 tok/s (2.98×)
  • c2 / c4: 33.3 / 55.2 vs 76.2 / 76.0 — Coder-Next scales, Laguna flattens
  • Resident footprint: 101 GiB vs 53.9 GiB (47 GiB reclaimed)
  • Cold start: 8–15 min vs ~90 s

Verdict

  • Keep coder: Qwen3-Coder-Next restored as the primary lane at 2 × 262,144.
  • Specialist: Laguna S 2.1 retained on demand for tool-calling and verified 257K retrieval. Swaps with coder, so only one large coding model is resident. (Retired as a named lane four days later, July 30.)
  • Promotion criteria corrected: devcode on this model spanned 83.6–92.1 — a ~9-point spread. Speed, scaling, and memory now carry the weight.

Ops after the revert

507 coder requests, 0 errors, 59.3 tok/s average, longest turn 123 s. Independent sparkrun comparison ranked Coder-Next first of ~30 Spark configurations (0.74 lm-eval mean, HumanEval 0.71 vs Laguna 0.37).

← All notes