Skip to content

The visual lab / A map of what we learned

See how the pieces work together.

Four machines, a changing model roster, and a notebook full of tests. Trace the requests, explore the tradeoffs, and follow the evidence that shaped the current setup.

01 / Current architecture

The request has a destination.

Full hardware & config
Inside the local labIllustrated flow / Oct 02
Coding + native vision routing through llama-swapQwen3.8-Flash-Next · NVFP4 routes to DGX Spark + GX10. The two GB10 hosts share an explicitly configured tensor-parallel coder. Every other route uses an individual host. Moving dots illustrate requests, not live activity.DGX SparkGB10 · 128 GBASUS GX10GB10 · 128 GBMac StudioM5 Max · 48 GBRTX 509032 GB GDDR7llama-swapROUTING + RESIDENCYTWO NODES · TP2 + EXPERT PARALLELISM

Coding + native vision

DGX Spark + GX10

Qwen3.8-Flash-Next · NVFP4
Persistent TP2 deployment

Current config; older-model retrieval results do not qualify this model. This illustration uses verified configuration, not live telemetry.

Verified 02 Oct 2026

Two GB10 hosts.
Two specialist machines.
One routing layer.

The DGX Spark coordinates the gateway and the primary coder. The ASUS GX10 supplies the second tensor-parallel worker. The Mac Studio handles fast agents, local coding, writing, and embeddings. The RTX 5090 takes media and perception workloads.

Select a role in the diagram to follow its configured route. Animated packets explain the architecture; they do not report current traffic.

Build your own baseline

Memory / Four separate budgets

Capacity is local.
Placement is deliberate.

The gateway selects a backend. Only the two GB10 hosts run an explicitly distributed coder. Routing the Studio and 5090 does not turn their memory into one model allocation.

Weights + KV cache + activations + runtime + headroom all share the usable budget. Quantization changes weights; context and concurrency change cache pressure.

Mac Studio

48 GB installed unified memory, 18 CPU cores. The September evaluation recorded a Metal recommended working set of 37.44 GiB; 48 GB is not an unrestricted model budget.

Budget & setup guide ↗

02 / The measured tradeoffs

A useful chart keeps the protocol.

All benchmark records

01 / Same Studio, different strengths

Speed is only one axis.

Protocol & limitations ↗

M5 Max · 40 GPU cores · 48 GB · Sep 22–23, 2026 · thinking off

HumanEval / % ↑

050100080160

Code decode / tok/s →

Coding fallback / MLX 4-bit

95.1%

HumanEval · 156 / 164 tasks passed

All 164 original tasks, one greedy chat sample, original executable tests, no repair or retries. Chat-adapted pass@1; training contamination is possible. Gemma was a writing-role evaluation and did not run this suite.

Speed: median of three short 512-output code probes on a loaded backend. Timing and quality use different prompts; this is not full-task wall time.

02 / Speculative decoding

One switch. Opposite outcomes.

Six draft positions helped code and hurt creative writing on the same Spark.

MTP offMTP on

Code

+57.4% throughput
64.1 tok/s
100.9 tok/s

Creative writing

-17.0% throughput
64.6 tok/s
53.6 tok/s

SCALE 0–120 tok/s

Tune speculation against the work you actually send. A faster coding probe does not establish a faster writing model.

Qwen3.6-35B-A3B Q6_K · llama.cpp b9967 · one quiet slot. Code temperature 0.3; creative 0.7. Repeat count and spread were not retained. The current TP2 coder uses a different model and three MTP positions.

Decision & source record ↗

03 / Serving more requests

More throughput. A different metric.

Four concurrent streams increase aggregate output. That total is shared across the requests.

One streamFour-stream aggregate

Lightning

77.6 tok/s
169 tok/s

Qwen 27B + DFlash2

19.6 tok/s
42.7 tok/s

Mistral Small 4

30.7 tok/s
75.9 tok/s

SCALE 0–180 tok/s

Track single-request latency and whole-server throughput separately when deciding how many slots to serve.

September 12 historical lab-pool rebaseline · vLLM 0.29 / SGLang. Aggregate values are not per-user speed. Physical placement is not specified in the retained table; these are not Studio or 5090 benchmarks.

Decision & source record ↗

04 / A promotion reversed

Keep the rollback easy.

A same-day comparison restored Coder-Next and demoted Laguna. The decision used time, quality, and memory.

Laguna SCoder-NextLOWER IS BETTER

July 26 comparison

435.7 s
82 s

SCALE 0–450 s

Coder-Next decoded 2.98× faster, completed the recorded task in 82 s versus 435.7 s, and reclaimed about 47 GiB. A model promotion is a reversible decision.

Historical same-node A/B · Laguna S 2.1 NVFP4 versus Qwen3-Coder-Next UD-Q4_K_XL. Custom criteria are not an autonomous success rate. Both are earlier models, not today's Flash-Next coder.

Decision & source record ↗

05 / Images have more than one budget

Fast, small, good: measure all three.

Image pipelines differ in warm latency and peak memory. Visual quality needs its own evaluation.

Recorded pipelineLOWER IS BETTER

Mage-Flow Turbo

3.96 s

Z-Image-Turbo

5.7 s

Sana Sprint

0.79 s

SCALE 0–6 s

Mage was the only tested pipeline to meet the recorded quality, sub-10 s, and under-20 GiB gate. Sana's lower latency did not make it the quality winner. Mage later moved on demand: a roughly 120 s cold start changes the experience.

July 22 historical lab-pool test · 1024² output. No numeric visual-quality score was retained. The October configuration places primary media on the 5090; that route audit does not re-measure these pipelines on that GPU.

Decision & source record ↗

06 / Speed did not pass the gate

A fast challenger stayed a challenger.

August's Lightning candidate decoded faster but missed the coding-quality gate.

Lightning candidateCoder-Next champion

August 11 comparison

70 %
86.4 %

SCALE 0–100 %

70% versus 86.4% was a 16.4 percentage-point gap, beyond the ±4-point noise band in the review. Lightning stayed on demand; it did not become the primary coder.

Historical model gate, not the later Studio MLX evaluation. Role-specific evidence matters: a candidate can miss the primary-coder gate and still be useful as a fast agent in a different deployment.

Decision & source record ↗

The evidence atlas / 22 records

Every study keeps its date.

Select a date to inspect its records. The dots count records, not performance.

MEASUREDOPERATIONSCATALOG

03 / Serving is part of the result

Residency, context, and useful latency.

Gateway & lifecycle guide

Serving policy / Illustrated, not live state

Keep the right model warm.

The Studio's primary swap group is exclusive. Embeddings have a separate persistent group. The supervisor owns the backend and proxy together so unloading releases both.

The same primary group also includes vision and image configurations. These buttons illustrate policy; they do not start any model.

Exclusive primary group and independent embeddings on the Mac StudioLightning is the selected illustration of the exclusive primary group. Embeddings remain in a separate persistent group. This describes configured lifecycle policy, not current residency.EXCLUSIVE PRIMARY GROUPLightningone large primary backendINDEPENDENT GROUPQwen Embedding · 4Bpersistent policy · Q8_0DISTINCT LIFECYCLES · SHARED HARDWARE

Fast agent

262K model ceiling; full capacity unverified · Exclusive primary swap group · TTL 600 s

TTL zero alone did not establish persistence. Passive metadata polling once woke inactive models. Availability, residency, and lifecycle are separate observations. Read the lifecycle guide ↗

Context / Settings and evidence

A long limit needs a long test.

Context limits, retrieval probes, concurrency, and cold prefill answer different questions. The current model inherits a configured budget; it does not inherit another model's qualification.

Measured / Jul 21

~241K

Earlier Coder-Next retrieval

Two deep sessions, warm affinity; 4.3 s warm turn. This does not establish two cold 256K requests at once.

Original qualification ↗

Measured / Sep 22–23

~34K

Studio synthetic retrieval

262K reference scale

Middle-depth synthetic checks. Qwen and Gemma have 64K serving budgets; Lightning's 262K ceiling remains unverified.

Protocol boundaries ↗

Configured / Oct 02

262,144

Current Flash-Next TP2 coder

8 configured slots · 8,192 batched tokens · 0.74 allocator. Eight full-context sessions were not certified.

Configuration audit ↗

SOLID = RETRIEVAL OBSERVATION · DASHED = CONFIGURED / REFERENCE LIMIT · DIFFERENT MODELS AND DATES

Jul → Oct 2026 / Learning in public

The setup changed. The lessons stayed.

All journal notes ↗

Chapter 01 / Jul 11, 2026

+57.4% code / −17.0% creative

Six-position MTP changed two workloads in opposite directions. The optimization became a per-workload decision.

Read the decision ↗