01Tools disappear through a proxy
First compare direct and routed responses. A proxy that copies only content can silently discard tool_calls. Preserve and test the response structure, including fragmented parallel calls, before blaming the model. The Studio evaluation demonstrated this failure directly.
02Unloading never releases the model
Check backend process ownership and memory after swapping. A shell that backgrounds the model then replaces itself with a proxy can escape the gateway’s lifecycle. Keep both processes under one supervisor and verify cancellation and cleanup.
03A context setting is larger than the verified result
Separate vendor maximum, runtime setting, allocated KV capacity and measured retrieval. Test multiple depths and edge positions with the exact artifact. A successful small retrieval probe does not certify a full window.
04Prompt-speed gauges look impossible
The July observability repair found synthetic timing data shadowing real llama.cpp measurements: negative prompt rates and values in the millions. Keep engine timings distinct from estimates and represent unavailable measurements explicitly. Validate units and cache accounting before trusting a chart.
05A faster model makes the workflow worse
Measure full task completion and correctness. Speculation helped our July code probe while slowing creative writing. A higher tool score or public rank also failed to predict several local coding promotions. Tune by role, using the same workload and a known-good control.
06New settings exhaust memory
Return to a qualified context and sequence count, reduce simultaneous residency, and inspect actual memory allocations. Treat a successful load, a short smoke test and a long concurrent soak as different evidence. Preserve the previous runtime and checkpoint until the replacement is qualified.