Pull requests / #1269

#1269 serve: restore a session file without holding its K/V in RAM

open · @pspranger-throw · 0 comentarios · En GitHub

Server & APINVIDIA / CUDAModels & quantsSecurity

Descripción

The existing RESTORE path reads the whole snapshot into host memory before it writes anything, so its peak is about the size of the file. A deep session is easily larger than the RAM left on a machine that is also serving the model (in the measurement below: a 198,830-token session is a 3.05 GiB file, with ~5 GiB free while the model is loaded), and the server answers 503 ("when the RAM to read the file is not there", DETAILS.md).

Shallow sessions restore fine, which makes the limit easy to miss.

This change restores in two passes over the same file, format v1 unchanged:

- The read pass parses and validates the whole file exactly as today (header, identity, size, geometry, every   count against this session's limits, payload hash) but does not materialize the K/V: the running state and the K/V headers are kept, the buffers stay empty, and each layer's five part sizes are reported to the caller.
- The apply pass re-reads the file and copies the K/V into the authoritative pools a 16 MiB block at a time
  (cudaMemcpyDefault, so a device pool and a device-mapped host pool both work). It is bound to the read pass: a file whose K/V layer count or any part size differs is refused before that layer or part is applied. The payload hash is recomputed and compared to the trailer at the end. The first applied block is the gate — before it a failure refuses cleanly (the live session is exactly as it was), after it the engine fail-stops rather than decode from bytes it cannot vouch for.
- The residency refill after the pool writes is kept per layer (kv_stream_reset / kv_ring_restore): without it,
  the modes that keep K/V in VRAM read stale slots while the pool bytes and the reuse count both look right.

Peak RAM is two 16 MiB block buffers and the running-state metadata, so a session of any depth restores under the same floor. The existing session_file_read is untouched: callers that hold the K/V in RAM behave exactly as before.

Measured on a machine with an RTX 3090 and an RTX 2060S, Qwen3.8-Flash-Next IQ3_S 3.5 bpw, int8 KV,
--kv-resident 32768, 262K context:

- a 198,830-token session (3.05 GiB file) restored in 3.4-3.5 s in a new engine process (read pass 1.6 s); the
  next turn reused 198,859 of 198,879 prompt tokens (99.99%) in 1.8 s, where the same turn after a cold read takes ~177 s;
- a 39,060-token session (833 MB file): saved in 959 ms, restored in 1.0 s, next turn 99.95% reused in 0.7 s;
- available RAM, sampled every 0.2 s on a 64 GB machine, never fell below 4.67 GiB while the model loaded and the session restored — the restore fits in the RAM the serving shape already leaves free. The read pass's RAM preflight (what the engine requires to be available before reading the file) is 856 MiB at the 262K context limits, and the K/V payload is not part of that number, so it does not grow with the session.

conversation_file_test: 207 checks pass — the 185 existing ones unchanged (format golden included) plus 22 new ones (block bounds, the read/apply size binding, a file changed between the passes refused, a grown or shrunk file refused before any block, truncated files, identity).

En el sitio

Enlaces a install, modelos, releases.