Pull requests / #1163

#1163 conversation cache: a re-parked conversation reserves at most 1/8 of each K/V buffer instead of up to 16 MiB, so fewer get evicted

open · @BlueKingMuch · 0 评论 · 在 GitHub 查看

NVIDIA / CUDAModels & quantsDocumentationWindows

描述

A conversation that was restored and is parked again keeps its unchanged K/V pages and appends the new ones (92a3d5f, from #189). The appended segment also reserves room for later turns, growing geometrically up to 16 MiB, and it does that in every K/V buffer of the image: k, v, their scales and the indexer rows of each QSA layer, and the draft layer's. With Qwen3.8-Flash-Next (IQ3_S, int8 K/V) that came to 0.4 to 0.5 GiB of unused capacity per conversation parked again, about half the image on top. The cache counts capacity against `--conversation-cache-mib`, so the reserve decides which conversations get evicted.

This bounds the reserve by an eighth of what the buffer held before the append (at least 64 KiB, and at most 16 MiB as before). Turns still fill the reserve instead of adding a restore transfer each: the 4096-append test keeps its bound, and 256 turns of 1 MiB onto a 16 MiB buffer take 28 segments instead of 17.

The docs said parking falls back to a full capture when the reserve would evict another conversation. 6cd29d9 removed that check, since a fresh capture never evicts less - which did not hold with this reserve: below, the images parked again are about 1.5x the fresh size. With the reserve bounded, the two differ by at most an eighth of the K/V, and the docs now say what the reserve is instead.

## Measured

On 0.1.40.1, this commit against v0.1.40.1. RTX 4080 SUPER 32 GB, Ryzen 7 5800X3D, 96 GB DDR4, Windows 11, IQ3_S, 262K context, `--conversation-cache-mib 2816 --conversation-cache-slots 3`, a fixed expert placement (`--adapt-every 1000000 --expert-cache 8031`), greedy, 64-token answers. A main agent P (30,000 tokens) and two subagents S1 and S2 with the same system prompt (15,000 each) take turns for two rounds, and every return adds 2,300 tokens: P1 S1 P2 S2 S1b S2b P3 S1c S2c P4.

Parked images from the engine log, against (1 + checkpoints) x 0.11 GiB + 14.9 KiB per token:

| conversation parked again | v0.1.40.1 | this PR | computed |
|---|---|---|---|
| P at S2 (32,429 tokens, 3 checkpoints) | 1.38 GiB | 0.92 GiB | 0.90 GiB |
| S2 at P3 (17,429 tokens, 3 checkpoints) | 1.08 GiB | 0.69 GiB | 0.69 GiB |
| P at S1c (34,791 tokens, 4 checkpoints) | 1.49 GiB | 1.09 GiB | 1.04 GiB |

What that does to the turns (prompt tokens read, time to first token):

| | v0.1.40.1 | this PR |
|---|---|---|
| S1c, subagent 1 in the second round | 17,242, 5.5 s (S1 evicted at P3) | 2,300, 2.9 s |
| S2c, subagent 2 in the second round | 17,242, 5.4 s | 2,369, 2.9 s |
| all ten requests | 116,147 read, 49.9 s | 86,332 read, 44.6 s |
| evicted | 3 | 1 |

(S2c in v0.1.40.1 loses S2 another way: S1c, reading from the shared system prompt, moves S2's checkpoint into the session and overwrites S2. That is the next PR: #1164)

With the cache on and the fixed placement, the main agent's four answers equal a run without subagents, in both runs. `conversation_cache_test` passes on Windows (its new checks fail on v0.1.40.1 without the change), as do `conversation_snapshot_test` (with CUDA) and `conversation_file_test`.

cc @jeremiahritchey - the reserve came with your incremental capture; I kept its shape and only bounded it.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。