贡献 / #1471

#1471 feat: optionally offload idle session state while keeping model weights loaded

open · @W1nge · 0 评论 · 去 GitHub 看

Server & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

说明

An idle serving process currently keeps its session VRAM and host checkpoint payloads allocated even when the model should stay loaded but no request is running. This adds opt-in session offload using the existing CUDA VMM implementation: commit a process-local snapshot, release physical session pages, then restore at the same virtual addresses on the next command.

The default remains disabled. The supported scope is one CUDA VMM GPU with fully resident KV; incompatible batch, pipeline, remote, multi-GPU and elastic-KV configurations are rejected. Writes finish before memory is released; write failure retains memory. Missing, corrupt or expired snapshots clear cached history and reread the next prompt. GPU remap failure is fatal. The server exposes cache state and released bytes in `/v1/status`. Usage and limitations are in `docs/IDLE_SESSION.md`.

Validation on Windows / RTX 2080 Ti, based on v0.1.40.3:

- Full engine build and native `idle_cache_test`: exact multi-region restore, CUDA graph replay after remapping, constants, corruption/truncation/missing/expired files and cleanup.
- 263 Python server tests passed; the cache-status test was rerun after final changes.
- Real IQ3_XXS, INT8 KV, MTP HTTP lifecycle passed twice: idle save and cache-hit restore, corrupt-file fallback, expiry fallback, cancellation followed by offload and successful generation. The 4096-context snapshot was 309,197,048 bytes, with approximately 5.3 s save and 367 ms restore in one local run. These are local operation timings, not a throughput claim.

Linux/HIP and other GPUs were not runtime-tested. Explicit conversation SAVE/RESTORE commands have not received a separate offloaded-state integration test. This is a focused port of our local idle-session implementation, adapted to upstream VmmRange and current checkpoint fields; no machine-specific tuning is enabled.

Prepared and tested with Codex at the repository owner's request.

本站相关内容

相关页面的快捷入口。