Pull requests / #960
#960 serve: prefix snapshots on disk - a new chat restores its system prompt instead of reading it again
closed · @konijiwa110 · 0 コメント · GitHub で見る
Server & APINVIDIA / CUDAModels & quantsDocumentationWindowsLinux
本文
## What
The system-prompt checkpoint from #62/#65 lives in the one K/V arena. It is gone as soon as a request with a
different start comes in (a title request, a subagent, a second client), and after every restart. Parking keeps it
inside the parked conversation, but lending it takes that conversation out whole, so the conversation reads its
history again on its next turn. With agent clients this means many new chats still start from token 0, and their
system prompt plus tool list is often 20-35K tokens.
`--prefix-cache-dir DIR` keeps that checkpoint in a file:
- **Save.** A request that read its system prompt has answered, so its `DONE` is already out. Its root checkpoint's
running state and the K/V up to that point, draft layer included, are streamed from the session into `DIR`. This
happens once per distinct system prompt.
- **Restore.** A later prompt that starts with a saved prefix restores it, then reads only the rest. It does this
when the live session or a batch slot cannot serve more of the prompt. A parked conversation of the same length
loses to the file, so it stays parked. The file stays on disk.
- **Disk and RAM limits.**
- `--prefix-cache-disk-mib N` (default 51200) caps the directory, removing the least recently used files first.
- `--prefix-cache-ram-mib N` (default 2048, 0 = disk only) also keeps recent snapshots in RAM, when
`--conversation-cache-min-free-mib` allows it.
- Without RAM, save and restore stream through one 64 MB buffer, so low memory does not block either one.
- **Stale files.** Each file records the engine build, the pack, `--native`, `--mtp`, `--kv` and `STRATA_BF16_TC`.
If any of these differs (a rebuilt engine counts), the file is deleted at start-up.
- **Validation.** Every restore goes through the same checks as a parked snapshot before anything is written.
- A restore that fails before writing keeps the session as it was.
- A restore that fails after it has started writing reads the whole prompt instead. It does not exit.
This works on one GPU without `--batch` only, and stays off unless the directory is set.
## Measured
Swift 1.5 IQ3_XXS, 256K context, `--kv-resident 32768`. Each prompt is a 30K-token system prompt plus a short user
message, with greedy decoding and 64 output tokens.
**RTX 3080 12 GB, i7-13700KF, 64 GB RAM, Windows 11, Samsung 980 PRO.** Settings: `--kv int8`,
`--prefix-cache-ram-mib 0`, `--adapt-every 100000`. This is the branch as submitted.
| request | reused | prefix restore | request time |
| --- | ---: | ---: | ---: |
| first chat, full read (30,046 tokens) | 0 | - | 26.19 s |
| same prompt after an engine restart | 30,028 | 656 ms (disk) | 2.54 s |
| new chat after a request with another start | 30,028 | 499 ms (disk) | 2.53 s |
- The answer restored after the restart (content and reasoning) is byte-identical to the full read's.
- Saving the 30,028-token prefix took 522 ms after `DONE`. The file is 577 MB.
**RTX 2080 Ti 22 GB, i5-12490F, 62 GB RAM, Ubuntu, NVMe.** Settings: `--kv k8v4`. The engine is 0.1.39 with this
change plus #711/#742/#743.
| request | reused | prefix restore | request time |
| --- | ---: | ---: | ---: |
| first chat, full read (30,044 tokens) | 0 | - | 26.10 s |
| new chat after a request with another start | 30,028 | 190 ms (disk) | 1.15 s |
- An 8.4K-token prefix restores in 63 ms from the RAM tier and in 218 ms from disk.
- Writing a file takes 213-424 ms.
- A K8V4 file is about 118 MB of running state plus 12.4 KB per token, so a 30K-token prefix is 490 MB.
I also checked whether the model still knows the restored prompt. I put a code word in the middle of an 8.4K-token
system prompt. After the restore, a newly worded question still got it right.
## Verification
- `prefix_store_test` (new, GPU) runs FP16, INT8 and Q4_0, each with resident and streamed K/V, plus K8V4. It saves
a prefix from a session that holds more than the prefix, then restores it three ways: streamed from the file, read
into RAM, and from RAM. After each restore, the K/V, indexer and running state are byte-identical to the state at
the prefix. It also checks four cases:
- the index survives a restart;
- a file with another identity is deleted;
- a disk budget for one file evicts the older file;
- a truncated file is dropped.
210 checks pass on Linux (sm_75) and on Windows (MSVC, sm_86).
- `conversation_snapshot_test` passes unchanged (2169 checks). The existing `conversation_kv_validate` and
`conversation_kv_restore` now share their size check and residency step with the new streaming functions.
## Docs
`docs/DETAILS.md` gets a new paragraph, "Prefix snapshots on disk", after the parked-snapshot notes. That section
now says parked snapshots do not survive a restart, but prefix snapshots do.
関連リンク
インストール・モデル・リリースへの站内リンク。