Pull requests / #1529

#1529 serve: the conversation disk tier writes in the background, keeps the root checkpoint, takes what RAM cannot hold (on #1480)

open · @konijiwa110 · 0 评论 · 在 GitHub 查看

Setup & installServer & APINVIDIA / CUDAModels & quantsWindowsLinux

描述

Builds on #1480 by @routhjim, which stacks #1271 by @ANBAL534 and #1269 by @pspranger-throw. Their commits are merged here unchanged (8cfd251), and my change is the one commit on top (daff417). If #1480 lands first, this rebases down to that one commit.

I had a similar disk tier running locally for a while. When I compared it with yours, the session-file base and the streamed save and restore were better than mine, so I moved the parts of mine that still helped onto your branch instead of opening a competing PR.

## What changes

- **Evictions are written in the background.** When the RAM cache evicts a conversation, the image goes to a writer thread (`spill_async`). Before, it was written and flushed inside the request that caused the eviction. The writer thread only calls `session_file_write`, and the index is only touched on the main thread: `wait()` joins the write and indexes the file. This happens before a disk match, before a streamed save, and at shutdown, so a file is never matched before it is complete. A failed write removes its file.
- **A spilled file keeps two checkpoints.** It keeps the deepest one, as SAVE does, and the shallowest one, the root of the chain. The root is usually the end of the system prompt, which is what a new chat or a subagent resumes from. The checkpoints in between only serve edits further back. Dropping them halves the file.
- **What RAM cannot hold goes to disk instead of being dropped.** With the RAM cache on, a conversation that cannot be parked (it is over the budget, or the physical-RAM admission refuses it) is streamed to the directory the way `--conversation-cache-disk-only` writes it. A disk hit larger than the RAM budget is streamed back the way that mode reads it. This applies on one GPU, without `--batch` or `--peer-device`, like disk-only.
- **Fixes along the way.**
  - `put()` called `make_room` without the spill callback. An image larger than its estimate could evict conversations that were then lost. They now go to the disk tier too, and so does an image `put()` refuses.
  - If the RAM admission refuses a park or a disk hit while a background write is still holding an evicted image, the request waits for that write and checks again.
- **`--conversation-cache-disk-min-tokens N`** (default 0, off). Conversations shorter than N are not written.
- **Free-space floor.** Every spill write keeps `--session-min-free-mib` free, as SAVE does.

With the defaults and the RAM cache off, behaviour is the same as #1480.

## Measured

All numbers come from this branch against #1480 as merged here (8cfd251), on `main` 6674a00.

**Setup.** RTX 2080 Ti 22 GB, i5-12490F, 62 GB RAM, Ubuntu, NVMe. Swift 1.5 IQ3_S, `--kv int8 --kv-resident 32768 --spec 4 --mtp`.

**Workload.** Three agent-style conversations take turns with `--conversation-cache-slots 1`. They share a system prompt; each first prompt is about 16K tokens, and each later turn adds about 2.2K. There are 3 rounds, an engine restart (SIGTERM), then 2 more rounds. Decoding is greedy with 24 output tokens. The table shows prompt-read time for the 12 later turns.

| | median | range | turns resumed | shutdown |
|---|---:|---:|---:|---:|
| #1480, `--conversation-cache-mib 4096` | 5.20 s | 4.72-6.12 s | 12 / 12 | 2.9 s |
| this PR, `--conversation-cache-mib 4096` | 4.51 s | 4.39-6.46 s | 12 / 12 | 4.2 s |
| #1480, `--conversation-cache-mib 256` | 14.83 s | 11.11-23.32 s | 0 / 12 | 3.0 s |
| this PR, `--conversation-cache-mib 256` | 5.01 s | 4.84-6.63 s | 12 / 12 | 2.9 s |

- **Background writes.** A write took 0.36-0.65 s, and a request waited 0-20 ms for it.
- **File size.** A spilled file was 603-698 MiB, against 1149-1617 MiB with every checkpoint.
- **256 MiB budget.** Nothing fits in RAM there. On #1480 those conversations are dropped. Here a streamed save took 0.37-0.47 s and a streamed restore 0.59-0.96 s (read pass 0.28-0.52 s).
- **Shutdown.** Every shutdown, including writing out the RAM cache, took under 5 s, so `serve/server.py`'s 10 s wait before it kills the engine was never reached.

**Tests.** `conversation_spill_test` has 34 new checks, and all 93 pass, also under ThreadSanitizer. The checks cover the background write, the index waiting for it, branching at the root, `put()` evicting through the callback, waiting in the destructor, and refusal on the free-space floor. The other conversation tests pass. I have built and tested this on Linux only, not yet on Windows.

## Not included

- **Mirroring every turn to disk (#1331).** A clean shutdown already writes the RAM cache out. A crash costs one prompt re-read, so I didn't think several times the IO was worth it.
- **No-MTP save and restore.** That is #1492. The disk tier still needs the MTP drafter loaded, like SAVE/RESTORE.

The discussion in #1331 (@shahrokhzargarpour) and #1489 (@CC-David-CC) also helped shape this. Thanks to all of you. Happy to rebase or split it differently if that suits the stack better.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。