Issues / #1034
#1034 vram_elastic / POST /v1/vram: response schema and failure semantics (Windows 11, RTX 4090)
open · @oliver-prog-0705 · 3 コメント · GitHub で見る
Setup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows
本文
Hi, and thanks for vram_elastic (#533) and for the detailed release notes. We are evaluating v0.1.39 and have not installed it yet. The setup is Windows 11 (WDDM) with an RTX 4090 24 GB. We want to keep one Strata server running, yield VRAM to another GPU workload for a while, and then take it back. We understand that vram_elastic shrinks the expert cache in segments and does not necessarily free all VRAM. For us, safe behaviour on failure matters more than speed. Could you clarify the points below? Pointers to the code or docs are fine as answers. 1. **Response schema.** What does `POST /v1/vram` return on success, on partial success (if that exists), and on failure? Which fields and which HTTP status codes? 2. **reserve_mib semantics.** If the requested amount cannot actually be made free, does the call fail, shrink partway, clamp, or report the amount actually freed? 3. **Reacquisition.** When reducing the reserve again (for example `reserve_mib: null` to return toward the startup reserve, or `reserve_mib: 0` to reclaim as much as possible), what happens if another process is still holding the VRAM Strata wants back? Does Strata remain at its current smaller expert-cache size, partially regrow, retry, block, return an error, or risk OOM? 4. **Resident allocations.** At maximum shrink, what stays in VRAM on purpose: dense weights, KV cache, CUDA context, anything else? Is there an expected minimum residual footprint? 5. **Crash and restart.** If Strata or the process that sent the request dies during or after a resize, is any elastic state persisted? On restart, does Strata simply rebuild from its configured startup state? 6. **Observability.** Does any API or status field report the current elastic state, the effective reserve, the expert-cache size, or the amount released? If not, is external VRAM monitoring the intended way to verify a resize? 7. **Windows/WDDM.** Is vram_elastic tested or supported on Windows 11 with NVIDIA/WDDM? Are there known TDR or stall concerns around shrink or regrow? We noticed #961 (a WDDM driver wedge during the verify window), and we are only asking whether you consider it relevant here, not suggesting vram_elastic causes it. 8. **Synchronization and idempotency.** Does the call return only after the resize has finished? Is repeating the same `reserve_mib` value safe? Thank you for your time.
関連リンク
インストール・モデル・リリースへの站内リンク。