Pull requests / #718
#718 serve: park conversations on disk, keyed by a client id; requests without one stay in RAM
closed · @evaanp · 0 コメント · GitHub で見る
Server & APIMulti-GPUNVIDIA / CUDAModels & quantsLinux
本文
Builds on #716 (parking with --layer-split); the first 12 commits are that PR. ## What Parking in RAM (#57) needs host RAM that a PC running this model has little of: the experts take it. A long agent conversation is 1-2.5 GiB of snapshot, so with a few sub-agents and side calls the budget is gone, and every switch back reads the whole prompt again. `--conversation-disk-dir DIR` parks on disk instead. - **Keyed by the client.** The header `X-Strata-Conversation: <id>` (or the body field `strata_conversation`) names the conversation; the engine gets it as `conv=`. The id only picks the files: the prompt must still continue what was parked, otherwise the request reads from the deepest checkpoint it shares. - **Delta parks.** A park writes the K/V pages from the first one the session rewrote, the running state and the checkpoints; pages the files already hold are kept. `--conversation-disk-gib N` (default 32) bounds the directory; the oldest conversation goes first. - **Requests without an id stay in RAM.** The RAM cache keeps serving them beside the disk: a side call parks and restores there, and never pushes a keyed conversation out. A session holding a keyed conversation is never parked in RAM too, and a keyed request is looked up only on disk. - **When a conversation does not continue its live end**, the engine logs where: `does not continue its live end: diverges at N of M (prompt P); live=...|... prompt=...|...`, with the token ids around it. That is how a client that re-renders its own history (a template, a tool-call parser, a tokenization the model did not sample) can be found; without the line the request just read from a checkpoint or from 0. - `tools/conversation_cache_parity.py`: `--scenario disk` (A with an id, B, A again, a request without an id, a rewind to a checkpoint) and `--scenario mixed` (A with an id, B and C without, interleaved), each against a run with no cache at all. ## Measured RTX 5060 Ti 16 GB + RTX 2000 Ada 16 GB (`layer_split 30`), Ryzen 7 9700X, 64 GB RAM, Samsung NVMe. On our own branch (v0.1.38 merged with this code plus unrelated patches of ours; the parking code is the same as here), IQ3_S, `--conversation-cache-mib 1024`: - `--scenario disk`: PASS (A/B/A, delta parks, a rewind to a checkpoint; tokens and main-model state byte-exact). ~2K-token conversation: the first park writes 254.5 MiB in ~0.36 s, a later one 113-226 MiB while keeping 141-255 MiB already on disk; coming back reads its running state in 0.2-0.4 s and restores 28.5 MiB in 50-75 ms. - `--scenario mixed`: PASS. Only A (with an id) parked on disk; B, B+, C and B++ (without one) parked in RAM, and B+ and B++ came back from RAM, each turn byte-exact against its conversation played through without a cache. Earlier, IQ2_XS with the 0.1.34-based version of this code: - 24K-token conversation: first park 0.7 GiB in 1.06 s, later parks 0.1-0.2 GiB; restore 355 MiB of K/V in ~0.27 s plus 0.4-1.1 s of running state. `--scenario disk`: tokens and main-model state byte-exact. - Behind an agent app with sub-agents (a day of use): a ~53K-token main conversation came back in ~1.4 s where reading it again took ~90 s; two sub-agents alternating per call switched in ~1.5 s each. Developed with an AI coding assistant; all numbers measured on the machine above (Linux, CUDA 13).
関連リンク
インストール・モデル・リリースへの站内リンク。