Pull requests / #1271
#1271 serve: --conversation-cache-spill-dir keeps evicted conversations on disk across restarts
open · @ANBAL534 · 0 コメント · GitHub で見る
Server & APIMulti-GPUNVIDIA / CUDA
本文
The RAM conversation cache (`--conversation-cache-mib`) parks whole conversations so a later request reuses their prefill instead of reading the prompt again. But a parked conversation is dropped the moment RAM pressure or `--conversation-cache-slots` evicts it, and every parked conversation is gone when the engine restarts. On a small-RAM PC the budget is tight, so the conversations that matter are the ones evicted; after a restart the long prompt is read again from token 0. `--conversation-cache-spill-dir DIR` keeps what the RAM cache evicts: the evicted conversation is written to DIR as an ordinary session file (the same format and model/config identity as the slot save/restore API) plus a small metadata sidecar, so a later request - or a restart - reads it back instead of re-reading the prompt. **TL;DR** - a parked conversation evicted from RAM is written to disk and can be resumed later, also after a restart - matching a disk conversation costs a small sidecar read, not a K/V read - a spilled file is an ordinary session file: one format, one identity rule, interchangeable with the slot save/restore API - opt-in and off by default; single-GPU, like the parking it extends ## What changes - A conversation the RAM cache evicts (`make_room`) is handed to a spill callback that writes it to the directory; on a clean `QUIT` or stdin close the parked conversations spill before exit (`spill_all`). - A new request matches the directory from the sidecars only (token and image lists, no K/V read) and, when a disk conversation offers a longer resume than RAM, the batch slots and the live state, reads it back whole - subject to the RAM budget and the `--conversation-cache-min-free-mib` floor. - `--conversation-cache-similarity F` and `--conversation-cache-n-min N` filter weak matches (least common-prefix fraction / token count); both default to 0, which keeps every exact prefix the RAM cache would have used. - The spill reuses `conversation_file.cpp` rather than adding a second on-disk format: a spilled conversation is byte-for-byte what the slot save/restore API would write, so the two are interchangeable and there is one format, one model/config identity rule, one checksum. ## Why it helps - **Less RAM, same reuse.** Set a small `--conversation-cache-mib` and let the disk tier hold the rest: the RAM cache stays a small hot working set, and the conversations it evicts stay resumable on NVMe instead of being recomputed. On a 12-16 GB card where the cache budget competes with the expert cache, this trades a little disk for prompt-read time. - **No cold load after a restart.** The engine logs `spill dir ready (N conversations, ...)` at start, and the first request after a restart resumes from disk instead of reading the whole prompt again. - **Prefill shared across conversations.** Because a spilled file is a session file, a hand-saved session and a spilled conversation are the same artifact; and similarity matching lets a new conversation that shares a long prefix with a spilled one (the same system prompt, the same document) resume from that prefix rather than from token 0. ## Flags | flag | default | meaning | |---|---|---| | `--conversation-cache-spill-dir DIR` | off | keep conversations the RAM cache evicts on disk, across restarts | | `--conversation-cache-disk-mib N` | 8192 | the directory's limit for this model (0 = off) | | `--conversation-cache-similarity F` | 0 | least common-prefix fraction a disk hit may offer (range [0,1)) | | `--conversation-cache-n-min N` | 0 | least common-prefix tokens a disk hit may offer | ## Testing - `conversation_spill_test` (41 checks, wired into `STRATA_BUILD_CONVERSATION_TESTS`): spill/match/load round-trip with the K/V compared byte for byte, similarity and n_min filtering, refusal of a foreign model/config identity, disk-budget eviction, reindex on reopen, and the RAM cache handing its evictions to the tier (`make_room` callback + `spill_all`). - The existing conversation cache / file / memory / split tests still pass. - `generate.cpp` and the module compile in the CUDA build and `strata` links; the flags appear in `--help` and bad values are refused. - A comparative end-to-end run (the same workload with and without the patch, measuring prompt-read time and RAM held) is still pending; I will add the measured numbers. ## Scope Single-GPU only: with `--layer-split` the disk tier stays off, matching the existing parking limit. A disk hit is read into host RAM before restore, so it must fit the RAM cache budget and leave the min-free floor available, or the request reads the prompt normally.
関連リンク
インストール・モデル・リリースへの站内リンク。