Skip to content

Journal

2026-07-30

SWE-bench leaderboard review

Lab study note from the internal benches record (page last-modified 2026-09-12). Not live Oct 1 telemetry.

The production coder sat 33rd on SWE-bench Verified (self-reported 70.6%). Of the 32 entries above it, only three both outranked it and fit one Spark at 256K native context. Every score on that board is self-reported (verified: false). A rank buys a gate slot, nothing more. Local numbers below were measured idle, thinking off.

Local gate versus Coder-Next

  • Ornith-1.0-35B — published 16 · 75.6%. Local: devcode median 84.2%, devtool 75.0%, c1 61.6 tok/s, 28.2 GiB.
  • Qwen3.6-35B-A3B — published 24 · 73.4%. Local: 81.7% / 75.0% / 62.4 tok/s / ~28 GiB.
  • Laguna-XS-2.1 — published 31 · 70.9%. Local: 83.9% / 50.0% / 35.3 tok/s / ~87 GiB.
  • Coder-Next incumbent — published 33 · 70.6%. Local: 86.4% / 58.4% / 50.9 tok/s / 53.9 GiB.

One gate run is not a result. Ornith’s five-run spread was 73.2, 81.7, 84.2, 85.7, 87.3 — median 84.2 against the champion’s 86.4 (median of 6), with roughly double the variance (stdev 5.6 vs 2.9). Swift stayed the hole: median 60.7 vs the champion’s 74.1.

Verdict — role split

All three rejected as primary coder. Ornith took the tools-specialist slot. coder-laguna was retired. Promoting on one lucky observation is how Laguna S 2.1 was promoted on July 23 and reverted three days later.

A -ub sweep on this architecture left 2048 as the standard: 4096 bought almost nothing at 30K and decode stayed flat.

← All notes