01Establish the platform
DGX Spark combines an Arm CPU and Blackwell GPU with 128 GB unified memory. Follow the vendor first-boot and update instructions, then record the driver and runtime. Use ARM64 builds whose kernels support your model and GB10. Our second GB10 system is an ASUS GX10, verified directly.
Read-only platform checks
uname -m
nvidia-smi --query-gpu=name,driver_version --format=csv
free -h02Qualify a single-node baseline
Verify the model, chat template, tool parser, thinking policy and retrieval before adding another node. On our GB10 machines, the GPU total-memory query reports N/A; use system memory and actual process/runtime allocations rather than substituting a fictional VRAM total.
03Separate routing from distributed inference
A gateway can select a host. Tensor parallelism splits model computation across workers. Multi-node vLLM needs a compatible execution environment, matching model artifacts and explicit distributed configuration; attaching two machines to a gateway does not implement that.
Our active configuration
# Inspected coder settings · 02 Oct 2026
# Inventory excerpt, not a standalone launch command
model: Qwen3.8-Flash-Next NVFP4
runtime: vLLM 0.30.0
nodes: 2
tensor_parallel_size: 2
expert_parallel: true
max_model_len: 262144
max_num_seqs: 8
max_num_batched_tokens: 8192
gpu_memory_utilization: 0.74
mtp_draft_positions: 304Treat configured context as a claim to test
The current coder has a 262,144 context setting and eight sequence slots. Neither value establishes that eight full-context jobs were qualified. Older Coder-Next retrieval results belong to that older model. Measure long prompts, concurrent admission and memory headroom on the exact runtime you intend to use.
Primary references
Upstream docs describe supported behavior. Lab-specific configurations and numeric results retain their inspection or experiment dates.