Pull requests / #62
#62 serve: the conversation cache keeps its shared prefix; the rest rotates LRU
closed · merged 2026-09-28 · @j-luwierski · 0 Kommentare · Auf GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quantsDocumentation
Beschreibung
## What The serve loop keeps up to six conversation checkpoints (~118 MB each, taken at turn boundaries and every `--prompt-cache-every` prompt tokens). The old policy was first-in-first-out: the oldest left — and the oldest is in practice the end of the system prompt, the prefix every new chat of the same client shares. Once it was gone, a new chat re-read the whole prefix at prefill speed. This changes the retention: the chain's root — the deepest point every kept conversation shares — is pinned, the rest rotates by least recent use, and mounting through a checkpoint re-dates it. `--prompt-cache 1` switches the pin off and returns the old behaviour exactly. ## Why a policy and not a radix tree The request loop erases any checkpoint that is not a prefix of the current prompt (the KV cache is one arena, one branch of history at a time), so the retained checkpoints are always a prefix chain — a radix cache's tree collapsed onto the one branch the session can hold. A trie over it would have exactly one path; the live question is only what to drop. `conv_cache.hpp` is that decision as one pure function, with the reasoning in its header comment. ## User-visible effect A new chat that shares a long system prompt starts reading after the prefix instead of from token 0. Measured on the first message of a new chat with a 16,747-token system prompt: **14,898 ms → 1,176 ms (12.7x)**; a 30K prompt saves roughly 25–30 s per chat start at ~1,100 tok/s. Coding agents and API clients that open a fresh conversation per task feel it on every task. Continuing an existing conversation, decode speed, the memory budget (still 6 × 118 MB) and — verified token for token — the answers themselves are unchanged. ## Verification - `conv_cache_test.cpp` — the policy's scenarios: the root is never a victim, leaves rotate LRU, a mount re-dates a checkpoint, the cap<2 guard. Passing (7 checks). - Engine A/B: IQ3_XXS on an RTX 4070 (sm_89), ten-request scenario, greedy, static expert residency (`--adapt-every 100000`); the driver and raw per-request metrics live outside the PR and can be attached on request. The measured request — the first message of a new chat sharing the system prompt: | run | RESUME | prompt read | | --- | ---: | ---: | | main (first-in-first-out) | 0 | 14,898 ms | | this branch | 16,384 | 1,176 ms | | this branch, repeat | 16,384 | 1,181 ms | | this branch, `--prompt-cache 1` | 0 | 14,948 ms | - All ten answers are token-for-token identical between the arms, the cache-serving one included (the restored read starts on a 2,048-token chunk boundary, so the fresh part's floating-point order matches a full read). - No regression elsewhere: the full first read 15.38 vs 15.36 s; decode tok/s equal per request (A1 48.4/48.4, B1 56.8/57.1, B2 63.2/63.2, A8 63.2/63.1). - Re-checked after rebasing on engine 0.1.18: RESUME 16,384, 1,183 ms. - `checkpoint_save`/`checkpoint_restore` and the state format are untouched, so `STRATA_CKPT_REREAD`'s token-for-token guarantee is unaffected. ## Docs `docs/DETAILS.md`: the conversation-cache paragraph and "Current limits" now describe the shared-prefix retention (switching chats re-reads only the part where they diverge).
Mehr auf der Site
Links zu Install, Modellen, Releases.