01Set up the host and verify the GPU
Install a supported Windows NVIDIA driver and WSL2 using the vendor instructions. CUDA in WSL uses the Windows driver; do not install a Linux NVIDIA display driver inside the guest. Our direct query found an RTX 5090, driver 617.14, and a WSL2 kernel.
WSL setup and verification
# In Windows PowerShell, if WSL is not installed
wsl --install -d Ubuntu
# In the WSL Linux shell
/usr/lib/wsl/lib/nvidia-smi02Build or select a CUDA backend
For a source build of llama.cpp, install its documented build prerequisites and a compatible CUDA toolkit, then enable the CUDA backend. Choose a reviewed source revision before compiling. Verify the startup log identifies the GPU.
llama.cpp CUDA build
# Run inside your reviewed llama.cpp checkout
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j 8
./build/bin/llama-server --help03Budget VRAM separately from host RAM
The GPU has 32 GB nominal GDDR7; the driver reported 32,607 MiB available as device capacity. CPU offload can make larger models load, but does not turn host RAM into equally fast GPU memory. Record offload settings, decode, first-response latency and full-task wall time.
04Schedule media and utility roles intentionally
The current gateway routes images, video, music, OCR, translation, transcription and speech to this workstation. Inventory also includes coding candidates. Large media jobs need an explicit admission policy so two pipelines do not overcommit VRAM.
Primary references
Upstream docs describe supported behavior. Lab-specific configurations and numeric results retain their inspection or experiment dates.