Pull requests / #62

#62 serve: the conversation cache keeps its shared prefix; the rest rotates LRU

closed · merged 2026-09-28 · @j-luwierski · 0 comments · View on GitHub

BenchmarksServer & APINVIDIA / CUDAModels & quantsDocumentation

Description

## What

The serve loop keeps up to six conversation checkpoints (~118 MB each, taken at turn boundaries and
every `--prompt-cache-every` prompt tokens). The old policy was first-in-first-out: the oldest left
— and the oldest is in practice the end of the system prompt, the prefix every new chat of the same
client shares. Once it was gone, a new chat re-read the whole prefix at prefill speed.

This changes the retention: the chain's root — the deepest point every kept conversation shares — is
pinned, the rest rotates by least recent use, and mounting through a checkpoint re-dates it.
`--prompt-cache 1` switches the pin off and returns the old behaviour exactly.

## Why a policy and not a radix tree

The request loop erases any checkpoint that is not a prefix of the current prompt (the KV cache is
one arena, one branch of history at a time), so the retained checkpoints are always a prefix chain —
a radix cache's tree collapsed onto the one branch the session can hold. A trie over it would have
exactly one path; the live question is only what to drop. `conv_cache.hpp` is that decision as one
pure function, with the reasoning in its header comment.

## User-visible effect

A new chat that shares a long system prompt starts reading after the prefix instead of from token 0.
Measured on the first message of a new chat with a 16,747-token system prompt: **14,898 ms → 1,176 ms
(12.7x)**; a 30K prompt saves roughly 25–30 s per chat start at ~1,100 tok/s. Coding agents and API
clients that open a fresh conversation per task feel it on every task. Continuing an existing
conversation, decode speed, the memory budget (still 6 × 118 MB) and — verified token for token —
the answers themselves are unchanged.

## Verification

- `conv_cache_test.cpp` — the policy's scenarios: the root is never a victim, leaves rotate LRU, a
  mount re-dates a checkpoint, the cap<2 guard. Passing (7 checks).
- Engine A/B: IQ3_XXS on an RTX 4070 (sm_89), ten-request scenario, greedy, static expert residency
  (`--adapt-every 100000`); the driver and raw per-request metrics live outside the PR and can be
  attached on request.

The measured request — the first message of a new chat sharing the system prompt:

| run | RESUME | prompt read |
| --- | ---: | ---: |
| main (first-in-first-out) | 0 | 14,898 ms |
| this branch | 16,384 | 1,176 ms |
| this branch, repeat | 16,384 | 1,181 ms |
| this branch, `--prompt-cache 1` | 0 | 14,948 ms |

- All ten answers are token-for-token identical between the arms, the cache-serving one included
  (the restored read starts on a 2,048-token chunk boundary, so the fresh part's floating-point
  order matches a full read).
- No regression elsewhere: the full first read 15.38 vs 15.36 s; decode tok/s equal per request
  (A1 48.4/48.4, B1 56.8/57.1, B2 63.2/63.2, A8 63.2/63.1).
- Re-checked after rebasing on engine 0.1.18: RESUME 16,384, 1,183 ms.
- `checkpoint_save`/`checkpoint_restore` and the state format are untouched, so
  `STRATA_CKPT_REREAD`'s token-for-token guarantee is unaffected.

## Docs

`docs/DETAILS.md`: the conversation-cache paragraph and "Current limits" now describe the
shared-prefix retention (switching chats re-reads only the part where they diverge).

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.