Issues / #533
#533 Hot VRAM resize: let the engine yield VRAM to other apps and take it back
closed · @AXKore · 3 commentaires · Sur GitHub
NVIDIA / CUDAModels & quantsWindowsLinux
Description
The engine sizes its VRAM budget once at start (expert cache + KV) and holds it for the whole run. On a desktop GPU shared with heavy apps (CAD, games), whichever starts second loses: WDDM evicts the engine's allocations to system RAM (decode collapses) or the other app fails. Issue #516 is the same pain on Linux; PR #380 adds a static WDDM budget; PR #378 makes K/V elastic inside the engine — but nothing yields VRAM to other processes yet.
Suggestion: an opt-in way to shrink the expert cache at runtime and hand the freed VRAM back, then reclaim it when it is free again. Even a coarse control would do — e.g. POST /vram {"reserve_gib": 8} on the API, or a simple half/full toggle, or STRATA_VRAM_FLOOR. The adaptive tier re-warms the returned experts quickly, and with expert_profile_save (#477) a shrink/re-grow cycle costs almost nothing.
Use case: an inference box that is also a workstation — run CAD or a game while inference keeps serving (slower, but functional) instead of fighting the driver for memory.
Environment: engine 0.1.36, Windows 11, RTX 4090 24 GB, IQ3_S 262K.Sur le site
Liens install, modèles, releases.