<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>pstuart.ai / Local AI journal</title><link>https://pstuart.ai/journal</link><description>A practical field guide and public lab journal for local LLMs: DGX Spark, RTX 5090, Mac Studio, measured benchmarks, setup guides, and agent workflows.</description><language>en</language><atom:link href="https://pstuart.ai/feed.xml" rel="self" type="application/rss+xml"/>
<item><title>Follow the config, not the old dashboard labels</title><link>https://pstuart.ai/journal/2026-10-02-configuration-audit</link><guid isPermaLink="true">https://pstuart.ai/journal/2026-10-02-configuration-audit</guid><pubDate>Fri, 02 Oct 2026 12:00:00 GMT</pubDate><description>Direct inspection corrected hardware identities and model placement: dual GB10 coding, Studio agents, and workstation utilities.</description><content:encoded><![CDATA[<p>The public site began with a dated October 1 snapshot and a static portal roster. A direct inspection of the machines, running launch arguments, active gateway configs and read-only dashboard endpoints showed that parts of that description were already stale.</p>
<h2>Hardware verified directly</h2>
<ul><li>NVIDIA DGX Spark and <strong>ASUS GX10</strong>, both GB10 unified-memory systems. The old site described the GX10 as a second DGX Spark.</li><li><strong>Mac Studio M5 Max</strong>, 18 CPU cores, 40 GPU cores, <strong>48 GB</strong> unified memory.</li><li><strong>RTX 5090</strong> workstation, <strong>Ryzen 9 9950X3D</strong>, Windows + WSL2. The GPU driver reports 32,607 MiB device capacity. Guest RAM is not presented as installed Windows RAM.</li></ul>
<p>Some dashboard hardware labels also lag the real machines. System reports take precedence for hardware facts; active configs take precedence for role placement.</p>
<h2>The current placement</h2>
<p>The primary coder is <strong>Qwen3.8-Flash-Next NVFP4</strong> on <strong>vLLM 0.30.0</strong>, across both GB10 nodes with tensor and expert parallelism. Its running launch arguments specify 262,144 context, eight sequences, an 8,192 prefill batch, allocator 0.74 and three native MTP draft positions.</p>
<p>The coordinator now routes <strong>agent and embeddings to the Studio</strong>. The Studio runs an exclusive primary swap group for its large models and image alternatives, with embeddings in a separate persistent group.</p>
<p>The workstation is the primary route for media and several utility roles: OCR, translation, transcription and speech. Several coding candidates also exist in its configured inventory. Configured does not mean benchmark-qualified or loaded.</p>
<h2>Keep the evidence separate</h2>
<p>No new inference benchmark was run during this read-only audit. These are configuration facts, not throughput claims. The older Spark studies and September Studio evaluation keep their own dates and protocols. Upstream deployment-kit changelog measurements were not republished as results from this lab.</p>
<p>The reusable lessons and current limits are collected in <a href="https://pstuart.ai/fleet" rel="noopener">hardware</a>, <a href="https://pstuart.ai/guides" rel="noopener">setup guides</a> and <a href="https://pstuart.ai/stack" rel="noopener">the software stack</a>.</p>]]></content:encoded></item>
<item><title>Snapshot, 1 October 2026</title><link>https://pstuart.ai/journal/snapshot-2026-10-01</link><guid isPermaLink="true">https://pstuart.ai/journal/snapshot-2026-10-01</guid><pubDate>Thu, 01 Oct 2026 12:00:00 GMT</pubDate><description>Public-safe facts from the 13:51:36 CDT lab pull. Dual-Spark TP2 coder-flash resident; live generation idle.</description><content:encoded><![CDATA[<p>Attributed <strong>2026-10-01 13:51:36 CDT</strong>. Dual Sparks had <code>coder-flash</code> / <code>coder</code> / <code>vision</code> loaded — Qwen3.8-Flash-Next NVFP4, <strong>262144</strong> context, TP2 across both GB10 boxes. About <strong>63 GiB</strong> GPU memory per Spark worker (~50% of a ~124.6 GiB unified pool). GPU util sat around <strong>92–96%</strong> from that residency.</p>
<p>Live generation was <strong>idle (0 tok/s)</strong> on every node. That is the only live throughput number from this pull. A later coordinator-page pull showed live gauges as connecting / incomplete, so this idle probe remains the published snapshot. No invented tok/s.</p>
<p>Catalog figures such as ~80–82 tok/s for agent / lightning, ~65 for agent-gemma, and ~60 for omni are dated roster copy; the Oct 1 probe saw agent unloaded. The coordinator roster (Capabilities) lists resident agent / lightning, embed, and Kokoro tts, with on-demand Mage-Flow image, omni, ocr, translate, stt, agent-gemma, tts-clone, and coder-ornith.</p>
<p>OCR, translate, STT, and TTS selectors were unloaded in the Oct 1 window. The Gaming PC (RTX 5090) peer was online and is the exclusive media box (music, video, image edit) — not the Spark chat worker. Mac Studio M5 Max had embeddings ready; its GPU block from that pull was mis-templated and is not published.</p>
<p>No private hostnames, no live charts, no public gateway URL, no connect recipes.</p>]]></content:encoded></item>
<item><title>The agentic repository improvement loop</title><link>https://pstuart.ai/journal/repo-prompt-loop</link><guid isPermaLink="true">https://pstuart.ai/journal/repo-prompt-loop</guid><pubDate>Thu, 01 Oct 2026 12:00:00 GMT</pubDate><description>A bounded pass, independent review, verified integration, and a durable record.</description><content:encoded><![CDATA[<p>The documented workflow uses a standing repo prompt to repeat useful, bounded improvement passes. Each pass inspects the repository, selects a concrete change, implements it, verifies it and creates a reviewable pull request.</p>
<h2>Review is a separate role</h2>
<p>QABot returns <strong>MERGE or HOLD</strong> from the evidence. MergeBot squash-merges eligible changes under repository policy. A HOLD sends the change back for correction and verification. Deployment authorization remains a separate policy decision.</p>
<h2>The reusable prompt</h2>
<p>The public <a href="https://pstuart.ai/workflow" rel="noopener">agent workflow</a> contains a reference prompt adapted from this process. It is a reusable technical template; private repository targets and original operating instructions are omitted.</p>
<p>The record for each pass should contain the problem, resulting behavior, tests actually run, review decision, limitations and next justified improvement. Project identities and business work stay private.</p>]]></content:encoded></item>
<item><title>The notebook is public. The machines are not.</title><link>https://pstuart.ai/journal/lab-notebook</link><guid isPermaLink="true">https://pstuart.ai/journal/lab-notebook</guid><pubDate>Thu, 01 Oct 2026 12:00:00 GMT</pubDate><description>First public note. What this site is, and what it will not publish.</description><content:encoded><![CDATA[<p>pstuart.ai is a lab notebook, not a product. The fleet is four local machines: two DGX Sparks (GB10), a Mac Studio M5 Max, and a gaming PC with an RTX 5090 under WSL. They pool through llama-swap — a private trusted-network gateway — for local coding and inference.</p>
<p>One Spark is the coordinator (resident agent / lightning, embeddings, Kokoro; on-demand Mage-Flow image and perception). The other is the GB10 coding peer (tensor-parallel Qwen3.8-Flash-Next, Coder-Next / long-context profiles, ornith, voice clone). The 5090 is exclusive media — music, video, image edit — not the Spark chat roster.</p>
<p>This site will not expose live internals, private hostnames, connect recipes, or a public API. The journal carries the study chronology from the internal benches record (page last-modified 2026-09-12): gates, rejects, and deploys, labeled as historical. The Oct 1 snapshot on the home page is a separate attributed pull — live generation that day was idle (0 tok/s). A later coordinator-page pull had incomplete live gauges, so that idle probe stays the published snapshot.</p>]]></content:encoded></item>
<item><title>The Studio gets its own measured roles</title><link>https://pstuart.ai/journal/2026-09-23-studio-evaluation</link><guid isPermaLink="true">https://pstuart.ai/journal/2026-09-23-studio-evaluation</guid><pubDate>Wed, 23 Sep 2026 12:00:00 GMT</pubDate><description>A 48 GB M5 Max evaluation separated fast agent, coder, and writer roles, and exposed proxy and lifecycle bugs.</description><content:encoded><![CDATA[<p>The September 22–23 evaluation ran sequential large-model tests on an M5 Max Mac Studio: 18 CPU cores, 40 GPU cores and 48 GB unified memory. This entry is a curated summary of the reviewed scorecard and numeric records, not a new October benchmark.</p>
<h2>Three models, three jobs</h2>
<p>Lightning MLX 4-bit was the fast-agent candidate. Qwen3.8-27B MLX 4-bit was the coder. Gemma-4 QAT Q4_K_M was evaluated as a writer through llama.cpp Metal.</p>
<ul><li>Chat-adapted HumanEval: Qwen <strong>156/164 (95.1%)</strong>, Lightning <strong>140/164 (85.4%)</strong>. Gemma was not run on this coding benchmark.</li><li>Code decode medians: Lightning <strong>150.1 tok/s</strong>, Qwen <strong>33.0 tok/s</strong>, Gemma <strong>78.1 tok/s</strong>. Each is a median of three short 512-output probes on a loaded backend.</li><li>Framework criteria: <strong>71.5% / 89.5% / 84.3%</strong>. These are regex checks, not compilation results.</li><li>Prose constraints: <strong>73.5% / 87.2% / 97.0%</strong>. Constraint adherence is not a literary-quality rating.</li></ul>
<h2>The serving path changed the answer</h2>
<p>The original metrics proxy reconstructed content and dropped non-streaming tool calls. Separate stochastic tool runs scored 32.4% through that proxy, 79.4% directly and 88.2% through the repaired path. The broken response structure was demonstrated; the difference between the latter two runs is not proof of a model quality gain.</p>
<p>The original launcher also left the MLX backend outside normal unload ownership. The repaired supervisor owns backend and proxy together. Embeddings gained an independent persistent group; TTL zero by itself had not kept them resident. Passive metadata polling was changed so it did not wake inactive models.</p>
<h2>What the measurements establish</h2>
<p>HumanEval used all 164 original tasks, one greedy chat sample, a 2048-output cap, thinking off, sandboxed executable tests and no answer repair. All reference solutions passed runner validation. This is chat-adapted pass@1, not EvalPlus, SWE-bench or a vendor leaderboard reproduction. Training-data contamination remains possible.</p>
<p>The local tools and prose suites are custom gates. Some developer-tool regex checks penalized valid alternatives, so their raw percentages are not autonomous coding success rates. The retrieved synthetic records reached about <strong>34K input positions</strong>, not a 64K or 262K capacity qualification. Manual Swift review found correctness issues even in answers that scored well.</p>
<p>The public <a href="https://pstuart.ai/benchmarks" rel="noopener">benchmark records</a> retain these distinctions.</p>]]></content:encoded></item>
<item><title>Production rebaseline + Mistral Small 4</title><link>https://pstuart.ai/journal/2026-09-12-production-rebaseline</link><guid isPermaLink="true">https://pstuart.ai/journal/2026-09-12-production-rebaseline</guid><pubDate>Sat, 12 Sep 2026 12:00:00 GMT</pubDate><description>Keep the split. Lightning owns latency, Qwen owns code, Mistral stays a quality watch — not a default.</description><content:encoded><![CDATA[<p>Lab study note from the internal benches record (page last-modified 2026-09-12). <strong>Not live Oct 1 telemetry.</strong></p>
<p>Both production lanes were remeasured after their runtime upgrades, then Mistral Small 4 119B-A6B NVFP4 ran the same fast gates on one Spark at 256K. Non-reasoning mode.</p>
<h2>Verdict — keep the split</h2>
<ul><li><strong>Keep agent.</strong> Lightning owns latency: 2.5× Mistral at c1 and 2.2× at c4. A long-form pass reached <strong>85.9 tok/s</strong> and <strong>0.13 s TTFT</strong>, though a requested 10K-word essay stopped at 3,935 words.</li><li><strong>Keep coder.</strong> Qwen3.8-27B + DFlash2 owns code: best framework score and <strong>94.1%</strong> broad tools, including perfect multi-turn follow-through and restraint.</li><li><strong>Retain checkpoint.</strong> Mistral wins broad tools (<strong>100%</strong>) and prose constraints (<strong>86.9%</strong>). It fits TP1 at 256K, but <strong>66.1 GiB</strong> weights and a ten-minute reload make it an unrouted specialist.</li></ul>
<h2>Fast-gate table</h2>
<ul><li>Framework code (8 tasks): Lightning <strong>80.3%</strong> · Qwen <strong>85.9%</strong> · Mistral <strong>81.4%</strong></li><li>Developer tool choice (12): <strong>75.0%</strong> · <strong>75.0%</strong> · <strong>83.3%</strong></li><li>Broad agent tools (34): <strong>82.4%</strong> · <strong>94.1%</strong> · <strong>100.0%</strong></li><li>c1 decode: <strong>77.6</strong> / <strong>19.6</strong> / <strong>30.7</strong> tok/s</li><li>c4 aggregate: <strong>169.0</strong> / <strong>42.7</strong> / <strong>75.9</strong> tok/s</li><li>Prose constraints: Lightning <strong>63.8%</strong> · Qwen not rerun · Mistral <strong>86.9%</strong></li><li>Ready time: Lightning and Qwen resident / instant; Mistral <strong>674 s</strong> cold · <strong>585 s</strong> cached</li></ul>
<h2>Capacity screen</h2>
<p>GLM-5.3-Flash NVFP4 is <strong>204.4 GB</strong> and DeepSeek-V4.1-Flash FP8 is <strong>510.3 GB</strong>. Neither fits one <strong>121.6 GiB</strong> Spark. Mistral’s one-Spark fit allocated <strong>66.12 GiB</strong> model memory plus a <strong>27.33 GiB</strong> KV pool; retrieval was not needle-tested in this fast gate.</p>]]></content:encoded></item>
<item><title>Nemotron-3.5-Lightning gate</title><link>https://pstuart.ai/journal/2026-08-11-lightning-gate</link><guid isPermaLink="true">https://pstuart.ai/journal/2026-08-11-lightning-gate</guid><pubDate>Tue, 11 Aug 2026 12:00:00 GMT</pubDate><description>Speed win, not a coder. Challenger kept as on-demand lightning. Qwen stays primary.</description><content:encoded><![CDATA[<p>Lab study note from the internal benches record (page last-modified 2026-09-12). <strong>Not live Oct 1 telemetry.</strong></p>
<p>Day-0 NVIDIA Nemotron-3.5-Lightning 30B-A3B (3B active), official NVFP4 + DSpark on a single DGX Spark. Question: is Lightning a coder upgrade, or a fast sub-agent / long-context challenger?</p>
<h2>Verdict</h2>
<ul><li><strong>Rejected as primary coder.</strong> Quality gap is decisive under the same rule that reverted Laguna: need a speed or capacity win without losing code. Lightning has both operational wins and still loses code by sixteen points.</li><li><strong>Kept as on-demand <code>lightning</code>.</strong> Sub-agent / multi-agent workhorse. Not in the coder spillover.</li><li><strong>Ornith stays the tools specialist.</strong> Lightning devtool <strong>66.7%</strong> beats the then-champion <strong>58.4%</strong> but loses to Ornith <strong>75.0%</strong>.</li></ul>
<h2>Measured (vs Coder-Next champion)</h2>
<ul><li>devcode: <strong>70.0%</strong> vs <strong>86.4%</strong> (−16.4, outside ±4 noise)</li><li>devcode wall: <strong>44 s</strong> vs <strong>85 s</strong> (1.93× faster)</li><li>Single-stream c1: <strong>82.3</strong> vs <strong>50.9</strong> tok/s (1.62×)</li><li>Aggregate c2 / c4: <strong>108.5 / 161.5</strong> vs <strong>76.2 / 76.0</strong> — Lightning scales; the champion plateaus</li><li>Footprint: ~<strong>22 GB</strong> + <strong>1.3 GB</strong> DSpark vs champion <strong>53.9 GiB</strong></li></ul>
<p>Speed alone does not promote a coder. Lightning earns a fleet seat as a challenger because the speed and capacity win is large and the quality loss is honest.</p>]]></content:encoded></item>
<item><title>Muse Glimmer dual-role reject</title><link>https://pstuart.ai/journal/2026-08-10-muse-glimmer-reject</link><guid isPermaLink="true">https://pstuart.ai/journal/2026-08-10-muse-glimmer-reject</guid><pubDate>Mon, 10 Aug 2026 12:00:00 GMT</pubDate><description>Meta’s open 30B rejected for both coder and always-on agent. Tools were real; decode was not.</description><content:encoded><![CDATA[<p>Lab study note from the internal benches record (page last-modified 2026-09-12). <strong>Not live Oct 1 telemetry.</strong></p>
<p>Day-0 open weights from Meta Superintelligence Labs (Apache 2.0). Dense multimodal 30B distilled from Muse Spark. Question: is it a fleet fit as coder or always-on agent?</p>
<h2>Verdict — rejected both roles</h2>
<ul><li><strong>Not a coding lane.</strong> devcode <strong>81.5%</strong> vs champion <strong>86.4%</strong> (−4.9, outside ±4 noise). devtool collapsed to <strong>25.0%</strong>. Wall-clock <strong>20×</strong> slower. Native context <strong>131K</strong> fails the verified-256K coder gate.</li><li><strong>Not always-on agent.</strong> Tools <strong>91.2%</strong> tie the Gemma agent, and developer sidecar <strong>30/30</strong> matches Omni. Dense 30B at <strong>10.1 tok/s</strong> loses to MoE Gemma at <strong>60.4 tok/s</strong> on every felt agent loop.</li><li><strong>Not justified as an on-demand specialist.</strong> Ornith already owns tools, Omni owns developer vision, Gemma owns fast agent.</li></ul>
<h2>Coder axes vs Coder-Next</h2>
<ul><li>c1: <strong>10.1</strong> vs <strong>50.9</strong> tok/s (0.20×). Even DFlash code peak <strong>34.3</strong> tok/s stays 0.67×.</li><li>GPU footprint ~<strong>21 GiB</strong> vs <strong>53.9 GiB</strong> — a win that does not pay for 0.2× decode.</li><li>Cold start <strong>17 s</strong> vs <strong>90 s</strong> — also not enough.</li></ul>
<p>Same promotion rule as the Laguna revert: do not promote on tool quality alone when single-stream is 0.2×. Not added to llama-swap.</p>]]></content:encoded></item>
<item><title>SWE-bench leaderboard review</title><link>https://pstuart.ai/journal/2026-07-30-swebench-review</link><guid isPermaLink="true">https://pstuart.ai/journal/2026-07-30-swebench-review</guid><pubDate>Thu, 30 Jul 2026 12:00:00 GMT</pubDate><description>Three models ranked above the coder. None won primary. Ornith took specialist; coder-laguna retired.</description><content:encoded><![CDATA[<p>Lab study note from the internal benches record (page last-modified 2026-09-12). <strong>Not live Oct 1 telemetry.</strong></p>
<p>The production coder sat <strong>33rd</strong> on SWE-bench Verified (self-reported <strong>70.6%</strong>). Of the 32 entries above it, only three both outranked it and fit one Spark at 256K native context. Every score on that board is self-reported (<code>verified: false</code>). A rank buys a gate slot, nothing more. Local numbers below were measured idle, thinking off.</p>
<h2>Local gate versus Coder-Next</h2>
<ul><li><strong>Ornith-1.0-35B</strong> — published 16 · 75.6%. Local: devcode median <strong>84.2%</strong>, devtool <strong>75.0%</strong>, c1 <strong>61.6</strong> tok/s, <strong>28.2 GiB</strong>.</li><li><strong>Qwen3.6-35B-A3B</strong> — published 24 · 73.4%. Local: <strong>81.7%</strong> / <strong>75.0%</strong> / <strong>62.4</strong> tok/s / ~28 GiB.</li><li><strong>Laguna-XS-2.1</strong> — published 31 · 70.9%. Local: <strong>83.9%</strong> / <strong>50.0%</strong> / <strong>35.3</strong> tok/s / ~87 GiB.</li><li><strong>Coder-Next incumbent</strong> — published 33 · 70.6%. Local: <strong>86.4%</strong> / <strong>58.4%</strong> / <strong>50.9</strong> tok/s / <strong>53.9 GiB</strong>.</li></ul>
<p>One gate run is not a result. Ornith’s five-run spread was 73.2, 81.7, 84.2, 85.7, 87.3 — median <strong>84.2</strong> against the champion’s <strong>86.4</strong> (median of 6), with roughly double the variance (stdev <strong>5.6</strong> vs <strong>2.9</strong>). Swift stayed the hole: median <strong>60.7</strong> vs the champion’s <strong>74.1</strong>.</p>
<h2>Verdict — role split</h2>
<p>All three rejected as primary coder. <strong>Ornith</strong> took the tools-specialist slot. <strong><code>coder-laguna</code> was retired.</strong> Promoting on one lucky observation is how Laguna S 2.1 was promoted on July 23 and reverted three days later.</p>
<p>A <code>-ub</code> sweep on this architecture left <strong>2048</strong> as the standard: 4096 bought almost nothing at 30K and decode stayed flat.</p>]]></content:encoded></item>
<item><title>Coder revert to Coder-Next</title><link>https://pstuart.ai/journal/2026-07-26-coder-revert</link><guid isPermaLink="true">https://pstuart.ai/journal/2026-07-26-coder-revert</guid><pubDate>Sun, 26 Jul 2026 12:00:00 GMT</pubDate><description>Same-node A/B reversed the July 23 Laguna promotion. Qwen restored; Laguna kept on demand.</description><content:encoded><![CDATA[<p>Lab study note from the internal benches record (page last-modified 2026-09-12). <strong>Not live Oct 1 telemetry.</strong></p>
<p>The first same-node, same-day A/B of the two coders, one resident at a time. The July 23 promotion of Laguna had been made on a single axis — code score — and every other axis favoured the incumbent. On this day the code score reversed too.</p>
<h2>Measured (Laguna S 2.1 vs Qwen3-Coder-Next)</h2>
<ul><li>devcode: <strong>87.2%</strong> vs <strong>90.9%</strong> (Coder-Next +3.7)</li><li>Same-suite wall: <strong>435.7 s</strong> vs <strong>82 s</strong> (5.3×)</li><li>devtool: <strong>75.0%</strong> vs <strong>66.7%</strong> — Laguna’s one durable edge</li><li>c1: <strong>17.7</strong> vs <strong>52.7</strong> tok/s (2.98×)</li><li>c2 / c4: <strong>33.3 / 55.2</strong> vs <strong>76.2 / 76.0</strong> — Coder-Next scales, Laguna flattens</li><li>Resident footprint: <strong>101 GiB</strong> vs <strong>53.9 GiB</strong> (<strong>47 GiB</strong> reclaimed)</li><li>Cold start: <strong>8–15 min</strong> vs ~<strong>90 s</strong></li></ul>
<h2>Verdict</h2>
<ul><li><strong>Keep coder:</strong> Qwen3-Coder-Next restored as the primary lane at 2 × 262,144.</li><li><strong>Specialist:</strong> Laguna S 2.1 retained on demand for tool-calling and verified 257K retrieval. Swaps with <code>coder</code>, so only one large coding model is resident. (Retired as a named lane four days later, July 30.)</li><li><strong>Promotion criteria corrected:</strong> devcode on this model spanned <strong>83.6–92.1</strong> — a ~9-point spread. Speed, scaling, and memory now carry the weight.</li></ul>
<h2>Ops after the revert</h2>
<p><strong>507</strong> coder requests, <strong>0</strong> errors, <strong>59.3 tok/s</strong> average, longest turn <strong>123 s</strong>. Independent sparkrun comparison ranked Coder-Next first of ~30 Spark configurations (0.74 lm-eval mean, HumanEval <strong>0.71</strong> vs Laguna <strong>0.37</strong>).</p>
<h2>Prompt-speed observability (same day)</h2>
<p>The dashboard had been recording coder prompt speed as unknown or worse: <strong>23 of 24</strong> stored samples were negative (min <strong>−121,862</strong>, max <strong>46,591,011</strong> tok/s), which flattened the prompt histogram into two of fourteen bins. Cause: a metrics shim synthesized timings for engines that omit llama.cpp’s timings block, then shadowed real llama.cpp measurements. After the shim left the coder path: <strong>0</strong> negatives, realistic band <strong>13–709</strong>, and an extra proxy hop on an identical 19-token prompt fell <strong>2,930 ms → 1,342 ms</strong>. The <strong>−1.0</strong> sentinel (not measured) is preserved. Not live Oct 1 telemetry.</p>]]></content:encoded></item>
<item><title>Laguna SM121 runtime</title><link>https://pstuart.ai/journal/2026-07-23-laguna-sm121</link><guid isPermaLink="true">https://pstuart.ai/journal/2026-07-23-laguna-sm121</guid><pubDate>Thu, 23 Jul 2026 12:00:00 GMT</pubDate><description>Serving tuned as far as it would go on GB10. Eager + .79 allocator shipped; DFlash K=7 rejected. Reversed July 26.</description><content:encoded><![CDATA[<p>Lab study note from the internal benches record (page last-modified 2026-09-12). <strong>Not live Oct 1 telemetry.</strong></p>
<p>Question: can Laguna’s serving runtime be tuned on GB10? Sweep covered DFlash, FP8 KV, CUDA graphs, and the allocator.</p>
<h2>Outcome</h2>
<ul><li><strong>Promoted (this day):</strong> eager mode with a <strong>.79</strong> allocator. <strong>.80</strong> crossed the earlyoom threshold during startup.</li><li><strong>Rejected:</strong> DFlash K=7. Draft accepted <strong>1,320 of 8,666</strong> tokens — <strong>15.2%</strong> — and reduced c4 throughput.</li><li>Best profile still only <strong>17.6</strong> tok/s c1 / <strong>54.6</strong> c4. KV pool allocated <strong>1,173,283</strong> tokens (<strong>4.48×</strong> the configured 262,144 window).</li></ul>
<p>The ceiling was the model on this hardware, not the configuration. This is the promotion the July 26 same-node A/B reversed.</p>]]></content:encoded></item>
<item><title>Image gen + Laguna / BTL-3 refresh</title><link>https://pstuart.ai/journal/2026-07-22-image-and-refresh</link><guid isPermaLink="true">https://pstuart.ai/journal/2026-07-22-image-and-refresh</guid><pubDate>Wed, 22 Jul 2026 12:00:00 GMT</pubDate><description>Mage-Flow Turbo 4B won the local 1024² envelope. The coder refresh produced a promotion that July 26 reversed.</description><content:encoded><![CDATA[<p>Lab study note from the internal benches record (page last-modified 2026-09-12). <strong>Not live Oct 1 telemetry.</strong></p>
<h2>Image generation — Mage-Flow Turbo 4B</h2>
<p>Six fixed prompts at 1024². Mage was the only tested model combining high visual quality, sub-10-second output, and under-20 GiB peak allocation.</p>
<ul><li><strong>Mage-Flow Turbo 4B</strong> — <strong>3.96 s</strong> warm, <strong>17.03 GiB</strong> peak. Production <code>image</code> alias. Not uncensored (life-drawing prompt returned a blank refusal).</li><li><strong>Z-Image-Turbo</strong> — <strong>5.70 s</strong> warm, <strong>21.69 GiB</strong>. Quality / text-rendering reference.</li><li><strong>Sana Sprint 1.6B</strong> — <strong>0.79 s</strong> warm, <strong>10.19 GiB</strong>. Preview specialist; weaker count fidelity.</li></ul>
<p><code>image</code> later moved on demand rather than resident — the pipeline had pinned <strong>17 GiB</strong> around the clock. Warm generation ~4 s; cold ~120 s.</p>
<h2>BTL-3 and Laguna refresh — superseded</h2>
<p>Fresh local runs of eight framework tasks and twelve tool-choice scenarios against refreshed NVFP4 snapshot 0761412.</p>
<ul><li>Laguna S 2.1: code <strong>89.0%</strong>, tools <strong>58.3%</strong>, c1 <strong>19.0</strong> tok/s, context edge <strong>257,107</strong>. About <strong>96.5 GB</strong>.</li><li>BTL-3: code <strong>88.0%</strong>, tools <strong>75.0%</strong>, c1 <strong>4.1</strong> tok/s. Always-on thinking exhausted its answer budget.</li><li>Then-current Coder-Next control: code <strong>83.6%</strong>, two-agent <strong>76.2</strong> tok/s.</li></ul>
<p><strong>Superseded.</strong> This study produced the July 23 Laguna promotion that the July 26 A/B reversed (Laguna then measured <strong>17.7</strong> tok/s single-stream).</p>]]></content:encoded></item>
<item><title>Three-lane coder and the July 21 gates</title><link>https://pstuart.ai/journal/2026-07-21-coder-lanes</link><guid isPermaLink="true">https://pstuart.ai/journal/2026-07-21-coder-lanes</guid><pubDate>Tue, 21 Jul 2026 12:00:00 GMT</pubDate><description>Three-lane dispatcher deployed; 2×256K qualified; Laguna S rejected; keep UD-Q4_K_XL; vision split to Nemotron Omni.</description><content:encoded><![CDATA[<p>Lab study note from the internal benches record (page last-modified 2026-09-12). <strong>Not live Oct 1 telemetry.</strong></p>
<p>Five controlled studies landed the same day.</p>
<h2>Three-lane coder — deployed</h2>
<p>Capacity-aware dispatcher. Two Spark slots plus one local lane. Four concurrent 256-token streams reached <strong>99.6</strong> aggregate tok/s (<strong>42.0–56.8</strong> per stream); excess queued, no 409s. Session affinity keeps a conversation on one worker so its prompt cache stays local.</p>
<h2>Coder 2 × 256K — qualified</h2>
<p>Will two independent <strong>262,144</strong>-token slots hold on one Spark? Yes. Both deep sessions retrieved correctly near <strong>241K</strong>. <strong>52.6</strong> tok/s single · <strong>76.0</strong> aggregate on two streams · warm turns <strong>4.3–4.4 s</strong>. About <strong>55.2 GiB</strong>. This is warm-session affinity, not a 2× cold accelerator.</p>
<h2>Laguna S vs coder — rejected</h2>
<p>First evaluation, including native TP=2. Code <strong>79.7%</strong> vs incumbent <strong>87.8%</strong>. Native TP=2 deadlocked on 512-token generations. Not a quality or throughput upgrade.</p>
<h2>Coder quantization — keep UD-Q4_K_XL</h2>
<p>Three Q5 variants beside production Unsloth dynamic Q4, same 256K settings.</p>
<ul><li><strong>UD-Q4_K_XL</strong> (keep): <strong>89.2%</strong> code, <strong>50.1</strong> tok/s, <strong>49.6 GB</strong></li><li>Q5_K_M: tools rose to <strong>66.7%</strong>, code fell to <strong>86.2%</strong></li><li>UD-Q5_K_M: <strong>87.4%</strong> code, ~<strong>17%</strong> slower</li><li>UD-Q5_K_XL: best Swift <strong>74.1%</strong>, still lost overall quality and speed</li></ul>
<p>More bits did not make a better default. Q5 cost <strong>7–10 GB</strong> and ~<strong>17%</strong> decode for no complete-workload win.</p>
<h2>Split vision router — Nemotron Omni</h2>
<p>Images are described once and replaced with cached text; coding stays on the text-only coder.</p>
<ul><li><strong>Keep vision:</strong> Nemotron Omni — <strong>30/30</strong> developer facts in <strong>10.0 s</strong>; production context raised to <strong>262,144</strong>. Perfect retrieval through <strong>240,123</strong> actual tokens.</li><li><strong>Keep coder:</strong> Qwen3-Coder-Next — <strong>87.6%</strong> code, <strong>75.0%</strong> developer tools, retrieval near <strong>255K</strong>.</li><li><strong>Fallback:</strong> Qwen3.5-9B — also 30/30, but <strong>80.7 s</strong> wall (8×).</li><li><strong>Does not fit:</strong> Nemotron Super 120B — loaded, then OOM-killed.</li></ul>]]></content:encoded></item>
<item><title>llama.cpp b10069 and the 256K coder bakeoff</title><link>https://pstuart.ai/journal/2026-07-20-runtime-and-coder</link><guid isPermaLink="true">https://pstuart.ai/journal/2026-07-20-runtime-and-coder</guid><pubDate>Mon, 20 Jul 2026 12:00:00 GMT</pubDate><description>Newer llama.cpp promoted on both Sparks. No challenger beat Coder-Next on framework code.</description><content:encoded><![CDATA[<p>Lab study note from the internal benches record (page last-modified 2026-09-12). <strong>Not live Oct 1 telemetry.</strong></p>
<h2>llama.cpp runtime — promoted</h2>
<p>Identical before-and-after tests on both Sparks. Production moved from <strong>b9967</strong> to <strong>b10069</strong> only after speed, cache, tools, long context, and multimodal gates passed. Active for agent, coder, embed, and omni. Prior image kept for rollback.</p>
<ul><li>Coder toolbench held <strong>94.1%</strong>. Retrieval and arithmetic remained 3/3 through <strong>121,512</strong> real prompt tokens.</li><li>Warm 17K prefix: <strong>16,668</strong> cached tokens, <strong>0.625 s</strong> warm vs <strong>13.63 s</strong> cold (<strong>21.8×</strong>).</li><li>External draft-simple no longer crashed but only <strong>30.2</strong> tok/s at 2K / <strong>13.6</strong> tok/s at 30K. Production kept ngram-mod.</li></ul>
<h2>256K coding + vision — keep Coder-Next</h2>
<p>Five native-256K challengers. Does any beat Coder-Next on framework code?</p>
<ul><li><strong>Keep Coder-Next:</strong> <strong>87.6%</strong>. Best challenger <strong>81.8%</strong>.</li><li>Nemotron Cascade 2 was the context-efficiency winner, not the code winner.</li><li>No challenger justified a production-route change.</li></ul>]]></content:encoded></item>
<item><title>Multimodal creative agent → Gemma-4-26B QAT</title><link>https://pstuart.ai/journal/2026-07-19-agent-gemma</link><guid isPermaLink="true">https://pstuart.ai/journal/2026-07-19-agent-gemma</guid><pubDate>Sun, 19 Jul 2026 12:00:00 GMT</pubDate><description>Four uncensored multimodal candidates. Gemma-4-26B QAT kept as always-on agent. Goetia and Fable rejected.</description><content:encoded><![CDATA[<p>Lab study note from the internal benches record (page last-modified 2026-09-12). <strong>Not live Oct 1 telemetry.</strong></p>
<p>Which model should be the always-on agent? Four candidates ran the same long-form, tool, vision, cache, concurrency, and 120K-context suite. The current Agent remained the strongest operational baseline.</p>
<h2>Verdict — keep Agent authoritative</h2>
<p><strong>Gemma-4-26B QAT</strong> won the job: best blind writing average, MoE-class latency, full context retrieval, no reliability regression.</p>
<ul><li>Blind overall: <strong>8.35</strong> / 10 (six constrained assignments)</li><li>Tools: <strong>91.2%</strong></li><li>Single stream: <strong>60.4</strong> tok/s</li><li>Four-way aggregate: <strong>129.9</strong> tok/s</li><li>Warm 17K prefix: <strong>0.303 s</strong> (<strong>23.2×</strong> over cold)</li><li>Deep edge pair: <strong>78.3 s</strong> · <strong>119,593</strong> prompt tokens</li></ul>
<h2>Rejected</h2>
<ul><li><strong>Goetia</strong> (Gemma-4 heretic): tools and vision matched; blind writing <strong>6.65</strong>. Operationally strong, not the writer.</li><li><strong>Qwen3.6 Fable Fusion:</strong> tools <strong>97.1%</strong>, but blind <strong>4.77</strong> and ~<strong>10.8</strong> tok/s. Rejected for the always-on seat.</li><li><strong>Ortenzya</strong> (dense 31B): blind <strong>8.17</strong>, interesting as a routed creative specialist at <strong>9.8</strong> tok/s — not the resident agent.</li></ul>
<p>Production restored with seven aliases live. This is dated study-page content, not the Oct 1 idle snapshot.</p>]]></content:encoded></item>
<item><title>Speculative decoding helped code and slowed writing</title><link>https://pstuart.ai/journal/2026-07-11-speculative-decoding</link><guid isPermaLink="true">https://pstuart.ai/journal/2026-07-11-speculative-decoding</guid><pubDate>Sat, 11 Jul 2026 12:00:00 GMT</pubDate><description>A quiet-box Qwen3.6 comparison showed why a drafting setting should be qualified per workload.</description><content:encoded><![CDATA[<p>This is a historical local measurement from the Spark tuning review, inspected again on October 2. <strong>Not live Oct 1 telemetry.</strong> External reference-rig numbers in that review are excluded from this public result.</p>
<h2>Same model, different acceptance</h2>
<p>Qwen3.6-35B-A3B Q6_K on llama.cpp b9967, one slot, quiet hardware. Native MTP used six draft positions.</p>
<ul><li>Code, temperature 0.3: <strong>64.1 → 100.9 tok/s</strong> with MTP enabled.</li><li>Creative writing, temperature 0.7: <strong>64.6 → 53.6 tok/s</strong> with MTP enabled.</li></ul>
<p>The local summary attributes the difference to draft acceptance: rejected speculative work can cost more than it saves. Repeat count and spread were not recorded in the reviewed summary, so this is not a universal speedup claim.</p>
<h2>Configure by role</h2>
<p>Keep speculation as a measured per-role choice. A result on this model and sampling policy does not transfer automatically to a different checkpoint, drafting method or runtime. The current Flash-Next MTP3 configuration is a separate record.</p>]]></content:encoded></item>
</channel></rss>
