Skip to content

Journal

2026-07-20

llama.cpp b10069 and the 256K coder bakeoff

Lab study note from the internal benches record (page last-modified 2026-09-12). Not live Oct 1 telemetry.

llama.cpp runtime — promoted

Identical before-and-after tests on both Sparks. Production moved from b9967 to b10069 only after speed, cache, tools, long context, and multimodal gates passed. Active for agent, coder, embed, and omni. Prior image kept for rollback.

  • Coder toolbench held 94.1%. Retrieval and arithmetic remained 3/3 through 121,512 real prompt tokens.
  • Warm 17K prefix: 16,668 cached tokens, 0.625 s warm vs 13.63 s cold (21.8×).
  • External draft-simple no longer crashed but only 30.2 tok/s at 2K / 13.6 tok/s at 30K. Production kept ngram-mod.

256K coding + vision — keep Coder-Next

Five native-256K challengers. Does any beat Coder-Next on framework code?

  • Keep Coder-Next: 87.6%. Best challenger 81.8%.
  • Nemotron Cascade 2 was the context-efficiency winner, not the code winner.
  • No challenger justified a production-route change.

← All notes