Skip to content

Journal

2026-09-12

Production rebaseline + Mistral Small 4

Lab study note from the internal benches record (page last-modified 2026-09-12). Not live Oct 1 telemetry.

Both production lanes were remeasured after their runtime upgrades, then Mistral Small 4 119B-A6B NVFP4 ran the same fast gates on one Spark at 256K. Non-reasoning mode.

Verdict — keep the split

  • Keep agent. Lightning owns latency: 2.5× Mistral at c1 and 2.2× at c4. A long-form pass reached 85.9 tok/s and 0.13 s TTFT, though a requested 10K-word essay stopped at 3,935 words.
  • Keep coder. Qwen3.8-27B + DFlash2 owns code: best framework score and 94.1% broad tools, including perfect multi-turn follow-through and restraint.
  • Retain checkpoint. Mistral wins broad tools (100%) and prose constraints (86.9%). It fits TP1 at 256K, but 66.1 GiB weights and a ten-minute reload make it an unrouted specialist.

Fast-gate table

  • Framework code (8 tasks): Lightning 80.3% · Qwen 85.9% · Mistral 81.4%
  • Developer tool choice (12): 75.0% · 75.0% · 83.3%
  • Broad agent tools (34): 82.4% · 94.1% · 100.0%
  • c1 decode: 77.6 / 19.6 / 30.7 tok/s
  • c4 aggregate: 169.0 / 42.7 / 75.9 tok/s
  • Prose constraints: Lightning 63.8% · Qwen not rerun · Mistral 86.9%
  • Ready time: Lightning and Qwen resident / instant; Mistral 674 s cold · 585 s cached

Capacity screen

GLM-5.3-Flash NVFP4 is 204.4 GB and DeepSeek-V4.1-Flash FP8 is 510.3 GB. Neither fits one 121.6 GiB Spark. Mistral’s one-Spark fit allocated 66.12 GiB model memory plus a 27.33 GiB KV pool; retrieval was not needle-tested in this fast gate.

← All notes