Pull requests / #8

#8 Conversation cache: keep the chat between requests, read only what is new

closed · merged 2026-09-26 · @Mirtraxxx · 0 comments · View on GitHub

BenchmarksServer & APINVIDIA / CUDAWindows

Description

## What it does

Right now every request starts from an empty session and reads its whole prompt again. In a long chat or an agent
loop, every turn waits for the full prefill. That's the main complaint in the r/LocalLLaMA thread ("throws away the
KV cache every prompt").

With this change the engine keeps the conversation and reads only the new part:

- **Live session.** If the new prompt starts with exactly the tokens the session already holds, it just continues.
- **Checkpoints.** Up to 6 are kept in RAM, about 118 MB each. One is taken at the last `<|im_start|>` of every
  prompt, where the history ends and the new assistant header begins, and one every 16K freshly read tokens. A
  rebuilt chat history (the client re-sends the messages, the template re-renders them) still matches the one
  taken before its last turn, so only the last reply and the new message are read.
- **Only the running state is copied.** That's the 36 GDN recurrences and conv histories, the 12 QSA indexer tails
  and the PLE history: the same set `Verifier::commit` rolls back. The KV cache and the pooled indexer keys are
  positional. A query only reads cells below it, and `dead` for its own block, so stale cells past a rewind point
  are never read and get overwritten.
- **Safety.** A checkpoint is used only if the prompt starts with exactly its tokens and pictures. Images are
  compared by a hash of their embeddings and grid, so a different picture of the same size never matches.
  Checkpoints whose cells the next request will overwrite are dropped.

## Measured

RTX 3090 24 GB, Ryzen 7 5700X, 64 GB DDR4-3000, Windows 11, Swift 1.5 IQ2_XS, `--expert-cache auto`, MTP `--spec 4`.

Agent-style session through `serve/server.py`: OpenAI streaming, one tool, thinking on, reasoning sent back each turn.

| Turn | Prompt | Reused | Read | Time to first token |
|---|---:|---:|---:|---:|
| 1 | 19,355 | 0 | 19,355 | 56.2 s |
| 2 (tool result added) | 20,882 | 19,407 | 1,475 | 5.6 s |
| 3 | 21,189 | 20,877 | 312 | 2.9 s |
| 4 | 21,355 | 21,184 | 171 | 2.6 s |

## Is it exact?

- `STRATA_CKPT_REREAD=1` makes the engine read a checkpoint's tokens again from position 0, in the same chunks the
  request that saved it used, instead of restoring it.
- `STRATA_STATE_HASH=1` prints a fingerprint of every part of the session after each request: GDN states, PLE
  history, indexer tails, pooled keys, the KV cells `[0, L)` of all 12 QSA layers, the MTP K/V, and `ple_prev`.
- A 4-turn chat (19K-token first prompt, two rebuilt follow-ups, one exact continuation) with `--adapt-swaps 0`, so
  the VRAM expert set is fixed. Restore and re-read gave **identical fingerprints for every component after every
  turn**, and identical tokens and draft counts.
- Two runs of the restore path were also identical to each other.

## Protocol (backward compatible)

- New stdout lines: `RESUME <n>` before reading, `PP <pos> <n> <ms> <fresh tok/s>` per chunk, and `REUSED <n>` once
  the prompt is read.
- `DONE` gains `<drafts accepted> <drafts offered> <reused>` after the existing fields.
- `serve/server.py` ignores lines it doesn't know. One change: `PP` lines now count as a heartbeat. They arrive
  every few seconds, which kept resetting the 10 s wait, so a long prompt sent no keep-alives at all: 0 in a
  41-second prefill before this change, 6 after.

## Options

- `--prompt-cache N`: checkpoints kept (default 6; `0` restores the old behaviour).
- `--prompt-cache-every N`: periodic checkpoint spacing (default 16384; `0` = only at the turn boundary).
- `--turn-token ID`: the token that opens a turn (default 248045, `<|im_start|>`).

## Limits

- One conversation is cached at a time. Two chats that take turns overwrite each other's state and re-read.
- Small reads still pay the prompt path's fixed cost (lending and refilling the borrowed cache slots): about 1–2.5 s
  even for a few dozen tokens.
- Only tested on Windows / one RTX 3090. Images through the cache are designed for but not tested end to end.

Made with Claude Code (Claude Opus 5.5); tested on the hardware above.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.