Measured / Jul 21
~241K
Earlier Coder-Next retrieval
Two deep sessions, warm affinity; 4.3 s warm turn. This does not establish two cold 256K requests at once.
Original qualification ↗The visual lab / A map of what we learned
Four machines, a changing model roster, and a notebook full of tests. Trace the requests, explore the tradeoffs, and follow the evidence that shaped the current setup.
01 / Current architecture
Coding + native vision
DGX Spark + GX10
Qwen3.8-Flash-Next · NVFP4
Persistent TP2 deployment
Current config; older-model retrieval results do not qualify this model. This illustration uses verified configuration, not live telemetry.
Verified 02 Oct 2026
The DGX Spark coordinates the gateway and the primary coder. The ASUS GX10 supplies the second tensor-parallel worker. The Mac Studio handles fast agents, local coding, writing, and embeddings. The RTX 5090 takes media and perception workloads.
Select a role in the diagram to follow its configured route. Animated packets explain the architecture; they do not report current traffic.
Build your own baselineMemory / Four separate budgets
The gateway selects a backend. Only the two GB10 hosts run an explicitly distributed coder. Routing the Studio and 5090 does not turn their memory into one model allocation.
Weights + KV cache + activations + runtime + headroom all share the usable budget. Quantization changes weights; context and concurrency change cache pressure.
48 GB installed unified memory, 18 CPU cores. The September evaluation recorded a Metal recommended working set of 37.44 GiB; 48 GB is not an unrestricted model budget.
Budget & setup guide ↗02 / The measured tradeoffs
01 / Same Studio, different strengths
M5 Max · 40 GPU cores · 48 GB · Sep 22–23, 2026 · thinking off
HumanEval / % ↑
Code decode / tok/s →
Coding fallback / MLX 4-bit
95.1%
HumanEval · 156 / 164 tasks passed
All 164 original tasks, one greedy chat sample, original executable tests, no repair or retries. Chat-adapted pass@1; training contamination is possible. Gemma was a writing-role evaluation and did not run this suite.
Speed: median of three short 512-output code probes on a loaded backend. Timing and quality use different prompts; this is not full-task wall time.
02 / Speculative decoding
Six draft positions helped code and hurt creative writing on the same Spark.
SCALE 0–120 tok/s
Tune speculation against the work you actually send. A faster coding probe does not establish a faster writing model.
Qwen3.6-35B-A3B Q6_K · llama.cpp b9967 · one quiet slot. Code temperature 0.3; creative 0.7. Repeat count and spread were not retained. The current TP2 coder uses a different model and three MTP positions.
Decision & source record ↗03 / Serving more requests
Four concurrent streams increase aggregate output. That total is shared across the requests.
SCALE 0–180 tok/s
Track single-request latency and whole-server throughput separately when deciding how many slots to serve.
September 12 historical lab-pool rebaseline · vLLM 0.29 / SGLang. Aggregate values are not per-user speed. Physical placement is not specified in the retained table; these are not Studio or 5090 benchmarks.
Decision & source record ↗04 / A promotion reversed
A same-day comparison restored Coder-Next and demoted Laguna. The decision used time, quality, and memory.
SCALE 0–450 s
Coder-Next decoded 2.98× faster, completed the recorded task in 82 s versus 435.7 s, and reclaimed about 47 GiB. A model promotion is a reversible decision.
Historical same-node A/B · Laguna S 2.1 NVFP4 versus Qwen3-Coder-Next UD-Q4_K_XL. Custom criteria are not an autonomous success rate. Both are earlier models, not today's Flash-Next coder.
Decision & source record ↗05 / Images have more than one budget
Image pipelines differ in warm latency and peak memory. Visual quality needs its own evaluation.
SCALE 0–6 s
Mage was the only tested pipeline to meet the recorded quality, sub-10 s, and under-20 GiB gate. Sana's lower latency did not make it the quality winner. Mage later moved on demand: a roughly 120 s cold start changes the experience.
July 22 historical lab-pool test · 1024² output. No numeric visual-quality score was retained. The October configuration places primary media on the 5090; that route audit does not re-measure these pipelines on that GPU.
Decision & source record ↗06 / Speed did not pass the gate
August's Lightning candidate decoded faster but missed the coding-quality gate.
SCALE 0–100 %
70% versus 86.4% was a 16.4 percentage-point gap, beyond the ±4-point noise band in the review. Lightning stayed on demand; it did not become the primary coder.
Historical model gate, not the later Studio MLX evaluation. Role-specific evidence matters: a candidate can miss the primary-coder gate and still be useful as a fast agent in a different deployment.
Decision & source record ↗The evidence atlas / 22 records
Select a date to inspect its records. The dots count records, not performance.
03 / Serving is part of the result
Serving policy / Illustrated, not live state
The Studio's primary swap group is exclusive. Embeddings have a separate persistent group. The supervisor owns the backend and proxy together so unloading releases both.
The same primary group also includes vision and image configurations. These buttons illustrate policy; they do not start any model.
Fast agent
262K model ceiling; full capacity unverified · Exclusive primary swap group · TTL 600 s
TTL zero alone did not establish persistence. Passive metadata polling once woke inactive models. Availability, residency, and lifecycle are separate observations. Read the lifecycle guide ↗
Context / Settings and evidence
Context limits, retrieval probes, concurrency, and cold prefill answer different questions. The current model inherits a configured budget; it does not inherit another model's qualification.
Measured / Jul 21
~241K
Two deep sessions, warm affinity; 4.3 s warm turn. This does not establish two cold 256K requests at once.
Original qualification ↗Measured / Sep 22–23
~34K
Middle-depth synthetic checks. Qwen and Gemma have 64K serving budgets; Lightning's 262K ceiling remains unverified.
Protocol boundaries ↗Configured / Oct 02
262,144
8 configured slots · 8,192 batched tokens · 0.74 allocator. Eight full-context sessions were not certified.
Configuration audit ↗SOLID = RETRIEVAL OBSERVATION · DASHED = CONFIGURED / REFERENCE LIMIT · DIFFERENT MODELS AND DATES
Jul → Oct 2026 / Learning in public
Chapter 01 / Jul 11, 2026
+57.4% code / −17.0% creative
Six-position MTP changed two workloads in opposite directions. The optimization became a per-workload decision.
Read the decision ↗