Pull requests / #1529
#1529 serve: the conversation disk tier writes in the background, keeps the root checkpoint, takes what RAM cannot hold (on #1480)
open · @konijiwa110 · 0 コメント · GitHub で見る
Setup & installServer & APINVIDIA / CUDAModels & quantsWindowsLinux
本文
Builds on #1480 by @routhjim, which stacks #1271 by @ANBAL534 and #1269 by @pspranger-throw. Their commits are merged here unchanged (8cfd251), and my change is the one commit on top (daff417). If #1480 lands first, this rebases down to that one commit. I had a similar disk tier running locally for a while. When I compared it with yours, the session-file base and the streamed save and restore were better than mine, so I moved the parts of mine that still helped onto your branch instead of opening a competing PR. ## What changes - **Evictions are written in the background.** When the RAM cache evicts a conversation, the image goes to a writer thread (`spill_async`). Before, it was written and flushed inside the request that caused the eviction. The writer thread only calls `session_file_write`, and the index is only touched on the main thread: `wait()` joins the write and indexes the file. This happens before a disk match, before a streamed save, and at shutdown, so a file is never matched before it is complete. A failed write removes its file. - **A spilled file keeps two checkpoints.** It keeps the deepest one, as SAVE does, and the shallowest one, the root of the chain. The root is usually the end of the system prompt, which is what a new chat or a subagent resumes from. The checkpoints in between only serve edits further back. Dropping them halves the file. - **What RAM cannot hold goes to disk instead of being dropped.** With the RAM cache on, a conversation that cannot be parked (it is over the budget, or the physical-RAM admission refuses it) is streamed to the directory the way `--conversation-cache-disk-only` writes it. A disk hit larger than the RAM budget is streamed back the way that mode reads it. This applies on one GPU, without `--batch` or `--peer-device`, like disk-only. - **Fixes along the way.** - `put()` called `make_room` without the spill callback. An image larger than its estimate could evict conversations that were then lost. They now go to the disk tier too, and so does an image `put()` refuses. - If the RAM admission refuses a park or a disk hit while a background write is still holding an evicted image, the request waits for that write and checks again. - **`--conversation-cache-disk-min-tokens N`** (default 0, off). Conversations shorter than N are not written. - **Free-space floor.** Every spill write keeps `--session-min-free-mib` free, as SAVE does. With the defaults and the RAM cache off, behaviour is the same as #1480. ## Measured All numbers come from this branch against #1480 as merged here (8cfd251), on `main` 6674a00. **Setup.** RTX 2080 Ti 22 GB, i5-12490F, 62 GB RAM, Ubuntu, NVMe. Swift 1.5 IQ3_S, `--kv int8 --kv-resident 32768 --spec 4 --mtp`. **Workload.** Three agent-style conversations take turns with `--conversation-cache-slots 1`. They share a system prompt; each first prompt is about 16K tokens, and each later turn adds about 2.2K. There are 3 rounds, an engine restart (SIGTERM), then 2 more rounds. Decoding is greedy with 24 output tokens. The table shows prompt-read time for the 12 later turns. | | median | range | turns resumed | shutdown | |---|---:|---:|---:|---:| | #1480, `--conversation-cache-mib 4096` | 5.20 s | 4.72-6.12 s | 12 / 12 | 2.9 s | | this PR, `--conversation-cache-mib 4096` | 4.51 s | 4.39-6.46 s | 12 / 12 | 4.2 s | | #1480, `--conversation-cache-mib 256` | 14.83 s | 11.11-23.32 s | 0 / 12 | 3.0 s | | this PR, `--conversation-cache-mib 256` | 5.01 s | 4.84-6.63 s | 12 / 12 | 2.9 s | - **Background writes.** A write took 0.36-0.65 s, and a request waited 0-20 ms for it. - **File size.** A spilled file was 603-698 MiB, against 1149-1617 MiB with every checkpoint. - **256 MiB budget.** Nothing fits in RAM there. On #1480 those conversations are dropped. Here a streamed save took 0.37-0.47 s and a streamed restore 0.59-0.96 s (read pass 0.28-0.52 s). - **Shutdown.** Every shutdown, including writing out the RAM cache, took under 5 s, so `serve/server.py`'s 10 s wait before it kills the engine was never reached. **Tests.** `conversation_spill_test` has 34 new checks, and all 93 pass, also under ThreadSanitizer. The checks cover the background write, the index waiting for it, branching at the root, `put()` evicting through the callback, waiting in the destructor, and refusal on the free-space floor. The other conversation tests pass. I have built and tested this on Linux only, not yet on Windows. ## Not included - **Mirroring every turn to disk (#1331).** A clean shutdown already writes the RAM cache out. A crash costs one prompt re-read, so I didn't think several times the IO was worth it. - **No-MTP save and restore.** That is #1492. The disk tier still needs the MTP drafter loaded, like SAVE/RESTORE. The discussion in #1331 (@shahrokhzargarpour) and #1489 (@CC-David-CC) also helped shape this. Thanks to all of you. Happy to rebase or split it differently if that suits the stack better. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
関連リンク
インストール・モデル・リリースへの站内リンク。