Pull requests / #1480
#1480 serve: --conversation-cache-disk-only - the conversation cache on disk, with no RAM budget (depends on #1271, #1269)
open · @routhjim · 0 comentarios · En GitHub
Server & APIAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentationWindowsLinux
Descripción
**Stacked on #1271 (spill directory) and #1269 (streaming restore).** The branch is #1271 (at 1a9e8b9) with #1269 (at d258f18) merged in. Only the last three commits are new: - dd146f4: the mode; - b189048: tests and docs; - ff30368: the cold-page-cache numbers in the docs. I will rebase once those two land, or fold this into #1271 if its author prefers. #1271 keeps the conversations the RAM cache evicts on disk. A conversation reaches disk only through a RAM eviction (`make_room`), so the RAM budget still decides what is kept. On a machine where the expert cache takes nearly all memory (Strix Halo: 72 GiB of experts in 128 GB), the conversation cache needs that budget most and can afford it least. `--conversation-cache-disk-only`, with `--conversation-cache-spill-dir`, keeps the conversation cache on disk only: - When a request switches to another conversation, the outgoing one is written to the directory the way a session SAVE writes it. The running state and the deepest checkpoint are copied; the K/V streams from its pools into the file (`conversation_snapshot_sources` + the streaming `session_file_write`). No host image is made. - When a conversation comes back, it is read the way #1269's RESTORE reads it. The read pass checks the whole file before anything on the GPU changes; a bad file is dropped, and the prompt is read as usual. Then the apply pass writes the pools 16 MiB at a time, with RESTORE's fail-stop after the first block. - `--conversation-cache-mib` is not used. `--conversation-cache-disk-mib` is the only limit, so no conversation is refused for its size. - The new file is written before this conversation's older copies are dropped. `drop_superseded` gains a `keep` path, because its rule alone would drop the new copy too. A failed write loses nothing. - The agent's next turn of the live conversation writes nothing. It resumes at the newest turn checkpoint, and the outgoing tail is only the reply the client sends again. - A clean shutdown saves the live conversation, flushed. Switch saves skip the flushes: a file torn by a crash fails the payload hash when it is read and is dropped. - Single GPU, without `--batch` or `--peer-device`; otherwise the flag says so and is off. **Why this matters beyond the budget.** With #1271 as it is, a conversation larger than the free RAM budget is neither parked nor spilled: ``` skip parking (snapshot 1292 MiB exceeds available budget) ``` If its disk copy was restored a turn earlier, that copy was already removed as superseded, so the conversation is gone, and its next turn reads 15,954 tokens from the start (18.9 s, in the run below). The long conversations are the ones that gain the most from a disk tier. Disk-only mode never asks the RAM budget. **Measured** on a Ryzen AI Max+ 395 (Radeon 8060S, gfx1151, 128 GB, Linux 7.2, ROCm 7.14.1, the iGPU alone). - Build: per `docs/STRIX_HALO.md`. - Model and flags: Qwen3.8-Flash-Next with K-quant experts, `--mmap-experts --expert-cache auto --spec 4 --mtp mtp/rt --lookup-chain 3 --kv int8 --max-context 65536`. Spill directory on the internal NVMe. - Workload: three agent-style conversations taking turns (A1 B1 C1 A2 ...). Each is a 9-12K-token first prompt, then turns that each add ~1.1K tokens of source text. Greedy, 64 output tokens. 3 rounds, an engine restart, then 2 more rounds. - Reuse and prompt time come from the `/metrics` totals around each request. | | prompt read, later turns (12) | turns resumed | |---|---:|---:| | no conversation cache | 11.4-22.3 s, nothing reused | 0 of 12 | | #1271, `--conversation-cache-mib 2048 --conversation-cache-slots 1` | 2.4-3.0 s, 88-94% reused | 11 of 12 (the one above lost) | | **this PR, `--conversation-cache-disk-only`** | **2.3-3.6 s, 88-94% reused** | **12 of 12** | - A switch wrote 360-443 MiB in 181-196 ms. A return read it back in 291-351 ms, of which the read pass was ~150 ms. - Cold disk: the same run again, with every spill file pushed out of the page cache before each request (`fsync`, then `posix_fadvise(DONTNEED)`; `fincore` showed 0 bytes resident each time). Restores took 275-361 ms (read pass 128-162 ms), and all 12 later turns resumed with the same 88-94% reuse. - The files are smaller than #1271's parked images (1.0-1.4 GB) because only the deepest checkpoint is saved, as SAVE already does. Three conversations took 1.3 GB on disk. - Process memory (`RssAnon`, sampled every 0.1 s) peaked 0.6 GiB above a run without the cache during a switch, and ended 0.15 GiB above it. No conversation stays in RAM. - Greedy replies matched the run without the cache for 13 of 15 turns. The other two (the first turn after the restart, and the turn after it) are identical to #1271's replies for the same turns: a restore and a full read of the prompt round differently. **Tests.** `conversation_spill_test`: 18 new checks, 59 pass. - A streamed spill writes byte for byte the file the image spill writes, matches and reads back with the same K/V. - It refuses a host K/V image or an empty conversation. - A failing source leaves no file and no index entry. - `drop_superseded` spares the `keep` copy, while the rule alone would drop it. The engine paths were run end to end as above. **Not tested:** - CUDA (no NVIDIA card here) and Windows; - `--kv-grow`: the restore calls `kvg_ensure` as the RAM restore does, but I did not run it with elastic K/V. Measured with the help of Claude (Anthropic). 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_013XBY17SSRKGmDXc2YFsrmp
En el sitio
Enlaces a install, modelos, releases.