Pull requests / #718

#718 serve: park conversations on disk, keyed by a client id; requests without one stay in RAM

closed · @evaanp · 0 commentaires · Sur GitHub

Server & APIMulti-GPUNVIDIA / CUDAModels & quantsLinux

Description

Builds on #716 (parking with --layer-split); the first 12 commits are that PR.

## What

Parking in RAM (#57) needs host RAM that a PC running this model has little of: the experts take it. A long agent
conversation is 1-2.5 GiB of snapshot, so with a few sub-agents and side calls the budget is gone, and every switch
back reads the whole prompt again. `--conversation-disk-dir DIR` parks on disk instead.

- **Keyed by the client.** The header `X-Strata-Conversation: <id>` (or the body field `strata_conversation`) names
  the conversation; the engine gets it as `conv=`. The id only picks the files: the prompt must still continue what
  was parked, otherwise the request reads from the deepest checkpoint it shares.
- **Delta parks.** A park writes the K/V pages from the first one the session rewrote, the running state and the
  checkpoints; pages the files already hold are kept. `--conversation-disk-gib N` (default 32) bounds the directory;
  the oldest conversation goes first.
- **Requests without an id stay in RAM.** The RAM cache keeps serving them beside the disk: a side call parks and
  restores there, and never pushes a keyed conversation out. A session holding a keyed conversation is never parked
  in RAM too, and a keyed request is looked up only on disk.
- **When a conversation does not continue its live end**, the engine logs where:
  `does not continue its live end: diverges at N of M (prompt P); live=...|... prompt=...|...`, with the token ids
  around it. That is how a client that re-renders its own history (a template, a tool-call parser, a tokenization
  the model did not sample) can be found; without the line the request just read from a checkpoint or from 0.
- `tools/conversation_cache_parity.py`: `--scenario disk` (A with an id, B, A again, a request without an id, a
  rewind to a checkpoint) and `--scenario mixed` (A with an id, B and C without, interleaved), each against a run
  with no cache at all.

## Measured

RTX 5060 Ti 16 GB + RTX 2000 Ada 16 GB (`layer_split 30`), Ryzen 7 9700X, 64 GB RAM, Samsung NVMe.

On our own branch (v0.1.38 merged with this code plus unrelated patches of ours; the parking code is the same as
here), IQ3_S, `--conversation-cache-mib 1024`:

- `--scenario disk`: PASS (A/B/A, delta parks, a rewind to a checkpoint; tokens and main-model state byte-exact).
  ~2K-token conversation: the first park writes 254.5 MiB in ~0.36 s, a later one 113-226 MiB while keeping 141-255
  MiB already on disk; coming back reads its running state in 0.2-0.4 s and restores 28.5 MiB in 50-75 ms.
- `--scenario mixed`: PASS. Only A (with an id) parked on disk; B, B+, C and B++ (without one) parked in RAM, and B+ and
  B++ came back from RAM, each turn byte-exact against its conversation played through without a cache.

Earlier, IQ2_XS with the 0.1.34-based version of this code:

- 24K-token conversation: first park 0.7 GiB in 1.06 s, later parks 0.1-0.2 GiB; restore 355 MiB of K/V in ~0.27 s
  plus 0.4-1.1 s of running state. `--scenario disk`: tokens and main-model state byte-exact.
- Behind an agent app with sub-agents (a day of use): a ~53K-token main conversation came back in ~1.4 s where reading
  it again took ~90 s; two sub-agents alternating per call switched in ~1.5 s each.

Developed with an AI coding assistant; all numbers measured on the machine above (Linux, CUDA 13).

Sur le site

Liens install, modèles, releases.