Issues / #533

#533 Hot VRAM resize: let the engine yield VRAM to other apps and take it back

closed · @AXKore · 3 Kommentare · Auf GitHub

NVIDIA / CUDAModels & quantsWindowsLinux

Beschreibung

The engine sizes its VRAM budget once at start (expert cache + KV) and holds it for the whole run. On a desktop GPU shared with heavy apps (CAD, games), whichever starts second loses: WDDM evicts the engine's allocations to system RAM (decode collapses) or the other app fails. Issue #516 is the same pain on Linux; PR #380 adds a static WDDM budget; PR #378 makes K/V elastic inside the engine — but nothing yields VRAM to other processes yet.

Suggestion: an opt-in way to shrink the expert cache at runtime and hand the freed VRAM back, then reclaim it when it is free again. Even a coarse control would do — e.g. POST /vram {"reserve_gib": 8} on the API, or a simple half/full toggle, or STRATA_VRAM_FLOOR. The adaptive tier re-warms the returned experts quickly, and with expert_profile_save (#477) a shrink/re-grow cycle costs almost nothing.

Use case: an inference box that is also a workstation — run CAD or a game while inference keeps serving (slower, but functional) instead of fighting the driver for memory.

Environment: engine 0.1.36, Windows 11, RTX 4090 24 GB, IQ3_S 262K.

Mehr auf der Site

Links zu Install, Modellen, Releases.