01Give every role a capacity and ownership policy
Our peer-aware llama-swap coordinator routes the coder to the dual GB10 backend, agents to the Studio and utility/media work to the workstation. A role name can stay the same when its host changes. Inspect the target and process owner instead of inferring placement from an alias.
02Make the resource group explicit
The following is a lifecycle excerpt from the active Studio configuration. It omits models, launch commands and endpoints. Match the schema to your exact gateway build; older upstream and peer-aware builds can differ.
Resource ownership, without private endpoints
# Lifecycle excerpt · current Studio config
routing:
router:
use: group
settings:
groups:
primary:
members: [lightning, qwen38, gemma-writer, agent-gemma, z-image, krea-2]
swap: true
exclusive: true
embeddings:
members: [embed]
swap: false
exclusive: false
persistent: true03Test the entire response contract
The original Studio metrics proxy dropped non-streaming tool calls while reconstructing text. Preserve tool-call IDs, names, arguments, streamed fragments, reasoning controls and usage fields. Compare direct and routed responses with deterministic fixtures; separate stochastic score movement from protocol correctness.
04Keep observation passive
Metadata polling previously woke models that should have remained unloaded. Passive metrics and inventory requests must not become implicit load requests. TTL controls idle unloading; persistence and exclusion decide what survives a competing model request.
05Expose a public record, keep the machines private
Publish model versions, limits, decisions and curated results. Keep credentials, network identities, user directories, private prompts and live control surfaces in the local deployment. This site has no connection to the inference gateway.
Primary references
Upstream docs describe supported behavior. Lab-specific configurations and numeric results retain their inspection or experiment dates.