Pull requests / #1271

#1271 serve: --conversation-cache-spill-dir keeps evicted conversations on disk across restarts

open · @ANBAL534 · 0 comments · View on GitHub

Server & APIMulti-GPUNVIDIA / CUDA

Description

The RAM conversation cache (`--conversation-cache-mib`) parks whole conversations so a later request reuses their prefill instead of reading the prompt again. But a parked conversation is dropped the moment RAM pressure or `--conversation-cache-slots` evicts it, and every parked conversation is gone when the engine restarts. On a small-RAM PC the budget is tight, so the conversations that matter are the ones evicted; after a restart the long prompt is read again from token 0.

`--conversation-cache-spill-dir DIR` keeps what the RAM cache evicts: the evicted conversation is written to DIR as an ordinary session file (the same format and model/config identity as the slot save/restore API) plus a small metadata sidecar, so a later request - or a restart - reads it back instead of re-reading the prompt.

**TL;DR**

- a parked conversation evicted from RAM is written to disk and can be resumed later, also after a restart
- matching a disk conversation costs a small sidecar read, not a K/V read
- a spilled file is an ordinary session file: one format, one identity rule, interchangeable with the slot save/restore API
- opt-in and off by default; single-GPU, like the parking it extends

## What changes

- A conversation the RAM cache evicts (`make_room`) is handed to a spill callback that writes it to the directory; on a clean `QUIT` or stdin close the parked conversations spill before exit (`spill_all`).
- A new request matches the directory from the sidecars only (token and image lists, no K/V read) and, when a disk conversation offers a longer resume than RAM, the batch slots and the live state, reads it back whole - subject to the RAM budget and the `--conversation-cache-min-free-mib` floor.
- `--conversation-cache-similarity F` and `--conversation-cache-n-min N` filter weak matches (least common-prefix fraction / token count); both default to 0, which keeps every exact prefix the RAM cache would have used.
- The spill reuses `conversation_file.cpp` rather than adding a second on-disk format: a spilled conversation is byte-for-byte what the slot save/restore API would write, so the two are interchangeable and there is one format, one model/config identity rule, one checksum.

## Why it helps

- **Less RAM, same reuse.** Set a small `--conversation-cache-mib` and let the disk tier hold the rest: the RAM cache stays a small hot working set, and the conversations it evicts stay resumable on NVMe instead of being recomputed. On a 12-16 GB card where the cache budget competes with the expert cache, this trades a little disk for prompt-read time.
- **No cold load after a restart.** The engine logs `spill dir ready (N conversations, ...)` at start, and the first request after a restart resumes from disk instead of reading the whole prompt again.
- **Prefill shared across conversations.** Because a spilled file is a session file, a hand-saved session and a spilled conversation are the same artifact; and similarity matching lets a new conversation that shares a long prefix with a spilled one (the same system prompt, the same document) resume from that prefix rather than from token 0.

## Flags

| flag | default | meaning |
|---|---|---|
| `--conversation-cache-spill-dir DIR` | off | keep conversations the RAM cache evicts on disk, across restarts |
| `--conversation-cache-disk-mib N` | 8192 | the directory's limit for this model (0 = off) |
| `--conversation-cache-similarity F` | 0 | least common-prefix fraction a disk hit may offer (range [0,1)) |
| `--conversation-cache-n-min N` | 0 | least common-prefix tokens a disk hit may offer |

## Testing

- `conversation_spill_test` (41 checks, wired into `STRATA_BUILD_CONVERSATION_TESTS`): spill/match/load round-trip with the K/V compared byte for byte, similarity and n_min filtering, refusal of a foreign model/config identity, disk-budget eviction, reindex on reopen, and the RAM cache handing its evictions to the tier (`make_room` callback + `spill_all`).
- The existing conversation cache / file / memory / split tests still pass.
- `generate.cpp` and the module compile in the CUDA build and `strata` links; the flags appear in `--help` and bad values are refused.
- A comparative end-to-end run (the same workload with and without the patch, measuring prompt-read time and RAM held) is still pending; I will add the measured numbers.

## Scope

Single-GPU only: with `--layer-split` the disk tier stays off, matching the existing parking limit. A disk hit is read into host RAM before restore, so it must fit the RAM cache budget and leave the min-free floor available, or the request reads the prompt normally.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.