Pull requests / #751
#751 Persist conversation KV state in a bounded disk LRU cache
closed · @Averyyy · 0 Kommentare · Auf GitHub
BenchmarksNVIDIA / CUDAModels & quantsSecurityWindows
Beschreibung
Completed conversations currently lose their cached state when the engine exits. This adds an opt-in disk LRU that preserves text, image and completed batch-slot state across graceful restarts, then restores the longest exact matching token/image prefix or turn checkpoint. Snapshots stream authoritative INT8 K/V, QSA/GDN/PLE state, image identity and prompt checkpoints for every GPU stage through 64 KiB buffers. Solo snapshots include MTP draft K/V. Batch snapshots use an explicit signature without draft state; restoring them clears the drafter's resident and host K/V while the main model verifies generated tokens. Completed slots commit before overwrite or shutdown, and pipelined groups drain and commit before padding overwrites a finished row. Cancelled and partial work is not committed. Slot-to-solo transfers retain the conversation's turn checkpoints. A request that finishes after leaving a slot can restore its original prompt or branch from disk after restart. The shared directory defaults to a 30 GiB budget with durable global LRU eviction, temporary writes and atomic replacement. Configure `--kv-persist`, `--kv-persist-dir`, `--kv-persist-max-mib`, `--kv-persist-identity` and `engine_close_s` through the existing server config. INT8, MTP and prompt checkpoints are required; RAM conversation parking stays disabled. The existing solo disk format remains compatible. Disk metrics accumulate across solo, admission, yielding and resumed request segments, alongside the upstream PCIe-offload metrics. Validation on Windows with two RTX 3080 20 GiB GPUs and Qwen3.8-Flash-Next IQ3_S: - Release/SM86 build, three CPU cache/persistence tests, and 274 server tests passed (five skipped). - Two-slot GPU tests passed slot overwrite, cancellation exclusion, graceful restart, exact batch continuation and solo MTP continuation. - A cancelled slot completed on the solo path preserved its user-turn checkpoint across restart and reproduced its reference output. - Four slots across two concurrently active pipeline groups passed 128-token output comparisons, restoration of every slot and solo continuation. - Vision tests passed embedding/grid identity changes, text switching, repeated-image no-write behavior, QUIT restoration, greedy output and valid primary-GPU state parity against a no-persistence baseline. - A 65,517-token document passed nine fact/arithmetic checks. After restart, the engine read 1.60 GB, restored 65,512 tokens and produced identical text. - The local deployment uses the latest upstream engine, native 262144 context and two parallel slots. Three concurrent document requests passed all 27 checks with thinking enabled, both before and after restart. Existing disk entries and local/public authentication were verified through the deployed service. The non-thinking document benchmark exposed one budget-arithmetic error (6976 instead of 7036), reproduced in a solo request; enabling thinking passed the same case.
Mehr auf der Site
Links zu Install, Modellen, Releases.