Issues / #1128
#1128 0.1.40 --kv-grow: a short request trims the K/V to 8192 cells and drops the conversation cache, so the next 128K+ continuation is re-read in full (20.9% of continuations, RTX 5090, IQ3_S)
open · @Darkstarrd-dev · 2 コメント · GitHub で見る
BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindows
本文
## Summary
With `--kv-grow`, a **short request trims the elastic K/V to 8192 cells** and the conversation cache goes with it, so the **next long continuation of a 128K+ conversation is re-read in full** (`prompt 128859 tokens = 0 reused + 128859 read in 28947 ms`).
On the same server and workload, the share of continuations that lose the cache entirely:
| engine run | `--kv-grow` | continuations | cache lost | rate |
| --- | --- | ---: | ---: | ---: |
| 0.1.28 | off | 20 | 0 | 0.0% |
| 0.1.29 | off | 212 | 1 | 0.5% |
| 0.1.36 | off | 412 | 3 | 0.7% |
| 0.1.38 | off | 34 | 0 | 0.0% |
| 0.1.39 | off | 567 | 8 | 1.4% |
| **0.1.40** | **on** | **43** | **9** | **20.9%** |
(continuation = a prompt ≥5000 tokens that equals the previous long prompt plus what it generated, within 1%; the numbers below give the 5% variant too)
So `--kv-grow` raises the full-re-read rate by roughly **15–20x** here.
## Environment
- Strata **0.1.40**, Windows 11 26200, engine from the release zip: `engine/BUILD.json` version `0.1.40`, `strata.exe` 146,444,800 B, sha256 `e41abc2efd0eb43d…`
- RTX 5090 D 32 GB (sm_120, driver 616.92), Intel Raptor Lake 8P+16E (no AVX-512, has AVX-VNNI), 64 GB RAM
- IQ3_S, `--max-context 262144`, `--kv` int8, **no `--kv-resident`** (the whole K/V in VRAM), images on, `--expert-cache auto`, `--prefill auto`, `--spec 4`, `--spec-min-p 0.7`, `--pcie-frac 0.20`, `--vram-reserve-mib 981`, **`--kv-grow`**
- Expert cache 11,189 slots (10,305–10,919 holding experts); the elastic K/V moves between 8,192 and 180,224 cells (0.66 → 2.34 GiB)
- Not about 0.1.40.1: the engine binary is the same one 0.1.40 shipped, and 0.1.40.1 only touches `serve/`
## Reproduction (plain engine log)
```
strata serve: prompt 127557 tokens = 127437 reused + 120 read in 670 ms (179.1 tok/s) <- long conversation, reused fine
strata: K/V trimmed to 8192 cells; 805 slots back to the expert cache, refilled from the profile
strata serve: prompt 204 tokens = 0 reused + 204 read in 882 ms <- a short request (new chat / title)
strata: K/V grown to 131072 cells (1.68 GiB); the expert cache gave 805 slots for it, 323 hotter experts moved ...
strata serve: prompt 128859 tokens = 0 reused + 128859 read in 28947 ms (4451.5 tok/s) <- the SAME conversation: full re-read
```
Every loss in that run has this shape: a `K/V trimmed to 8192 cells` triggered by a 53/107/124/172/204/251-token request, then a grow, then the long prompt with `0 reused`.
The nine losses (band 1%), with what the continuation carried:
| P | prefill | the previous long prompt |
| ---: | ---: | ---: |
| 108,610 | 23.6 s | 108,040 + 5312 |
| 110,837 | 25.2 s | 108,610 + 2177 |
| 113,456 | 25.3 s | 110,837 + 154 |
| 128,859 | 28.9 s | 127,557 + 561 |
| 129,742 | 29.4 s | 128,985 + 542 |
| 128,229 | 28.9 s | 124,858 + 3151 |
| 131,430 | 29.0 s | 128,229 + 3151 |
| 133,023 | 31.1 s | 131,430 + 102 |
| 157,034 | 34.9 s | 155,651 + 1599 |
In the same run the cached continuations take 0.4–2.5 s of prefill, so each loss costs about **25–35 s**.
## Why
`kvg_trim` (`src/program/generate.cpp:5362`) and `kvg_ensure` (`:5258`) take only a target cell count. They never look at the conversation checkpoints, and the elastic K/V shares one VMM range with the expert cache, so each shrink/grow remaps chunks and moves slots.
The trim is called at `:8405` with `n + 256`, where `n` is the **current (short) prompt's** length:
```cpp
// the elastic K/V: the outgoing session is parked and this request rewrites every cell from `resume` on,
// so cells past this prompt's are no longer anyone's - far more than it needs go back to the cache
if (kvg.on) { kv_quiesce(); if (!kvg_trim(n + 256)) { ... } }
```
For a 204-token request that means "keep 8192 cells", even though the conversation cache holds a 127,557-token checkpoint whose K/V lives in those cells. Right after it, the checkpoint erase at `:8398` and `live_ok = false` at `:8400` drop every checkpoint past the new (tiny) prompt:
```cpp
checks.erase(std::remove_if(checks.begin(), checks.end(), [&](const ConvCheckpoint& c) {
return (int64_t) c.ids.size() > resume || !starts_with(c.ids, c.imgs);
}), checks.end());
```
The next long prompt then finds `resume == 0` and reads all 128K tokens again. The comment's "the outgoing session is parked" does hold for a short request that directly follows a long one — `:8376` (`resume = parked.tokens`) restores its token count — but after the trim to 8192 cells there is no cell left to restore a 128K prefix from, so the parked/checkpointed prefix is gone too.
## Suggested fix
Make the trim conversation-aware, e.g.:
- keep at least `max(cells this request needs, deepest live/checkpointed prefix + its turn)` cells, or
- skip the trim while `checks`/`live` hold a prefix longer than the new prompt, or
- let `--kv-grow` grow without ever shrinking (a `STRATA_KV_GROW_TRIM=0`-style opt-out). There is no such flag today; `:8405` is the only `kvg_trim` call site.
## Cost/benefit on this machine
- `--kv-grow` gain, same day, same desktop, slot-matched A/B (10,333–10,397 → 11,990–12,101 slots): decode **123.7 → 135.2 tok/s (+9.3%)**
- `--kv-grow` cost in the run above: **281 s of wasted prefill** in 193 requests, against about **83 s** saved by the faster decode in the same run
For long conversations it is a net loss here, so I have turned `--kv-grow` off (decode falls to ~123.7 tok/s, slots to ~10,333). At the 5% band the same run is 16 losses / 504 s (30% of its total wall time) — the conclusion does not depend on which band is used.
I can attach the full server log or a trimmed excerpt — tell me which you prefer.
関連リンク
インストール・モデル・リリースへの站内リンク。