Issues / #661
#661 Proposal: save and restore a conversation to disk (/slots/0?action=save|restore)
closed · @maverde73 · 3 commentaires · Sur GitHub
Setup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsSecurityWindowsLinux
Description
## Proposal
Let the engine write the conversation it holds (what a parked conversation snapshot already holds) to one file and
read it back later, also after a restart of the same engine version with the same model and settings. This
implements llama-server-compatible save and restore requests and response fields for Strata's single slot, slot 0;
the feature is opt-in through `--slot-save-path DIR` (or `"slot_save_path"` in the config):
```
POST /slots/0?action=save {"filename": "chat1.bin"}
POST /slots/0?action=restore {"filename": "chat1.bin"}
```
It is not complete `/slots` support (no erase, no other slots) and the file format is Strata's own, not llama.cpp's.
Underneath, `strata --serve` takes two new stdin lines between requests: `SAVE <path>` and `RESTORE <path>`.
## Why
With a long conversation the first turn after a restart reads the whole prompt again: about 26 s at 63K tokens on
an RTX 4070 Ti. After the model is loaded, restoring the saved state took 0.97 s in the run below, and the next
32-token turn took 1.14 s. Engine startup itself still takes about 70-95 s. Useful for:
- agents and coding assistants that come back to the same long context (repository, documents) across restarts;
- switching between several long chats beyond the in-RAM conversation cache;
- unloading the model (`POST /unload`, a scheduled restart) without losing the context.
Files require the same engine version, model inputs and compatible settings: a file written by another engine
version is refused, not misread (same-version rebuilds are not identified by a build hash). It is the same idea as v0.1.36's `--expert-profile-save` (#477),
which keeps the learned expert profile across restarts; this keeps the conversation.
## Measured
RTX 4070 Ti 12 GB, Ryzen 9 5900X, 64 GB RAM, NVMe (ext4), Linux, v0.1.38, IQ3_XXS with the setup's config, a
62,993-token prompt + 32 generated tokens (63,025 tokens held), `--max-context 65536`:
| | |
| --- | --- |
| file size | 1.20 GB |
| save, flushed to disk | 0.87 s; 10 more saves (5 replacing one file, 5 new) 0.65-0.82 s, engine-reported |
| restore (new process, model loaded) | 0.97 s |
| next turn after restore (32 tokens) | 1.14 s |
| same turn without saving/restarting | 1.00 s |
| first turn without a session file (cold) | 25.6 s |
The turn after the restore produced the same 32 token IDs as the same next turn without a restart (this compares
two warm paths, not a claim of equality with one uninterrupted generation). A symbolic link, a second hard link, a
file with one flipped byte and a file restored by an engine with another RoPE base were refused, and the engine kept
serving after each refusal.
This saves conversation state, not all process-local execution history, so exact token replay across restarts is
not guaranteed: expert residency and CPU/GPU rounding can change later output. In the same 63K test the first
32-token continuation matched the process that stayed running; on the next continuation three restored processes
agreed with one another but differed from that process from the 28th token on. A refused corrupted restore in
between did not change the restored continuation. The cause of this divergence has not been isolated.
With the feature unused there was no measurable prefill, decode or peak-VRAM regression
in three alternating pairs at 3.4K, 29.6K and 63K, with identical output token IDs in all nine pairs.
## Design (short)
- One file: running state, the deepest checkpoint (the next turn's resume point), every QSA layer's K/V up to the
conversation's length and the draft layer's K/V.
- Format v1, little-endian: 64-byte header (magic, version, model fingerprint, config fingerprint, payload length,
header hash), payload, payload hash, end marker. Written to a new, exclusively created temporary file, flushed,
renamed over the target and the folder flushed (POSIX). A failure before the rename keeps the old file; a failed
folder flush after it is reported as such (the new file is already in place).
- Restore opens the file without following links and accepts only a regular file with one name. File and runtime
validation (size, header, both fingerprints, a RAM preflight, geometry and layer range, every count against the
bytes left and the session's exact limits, the
payload hash, the snapshot validation) complete before any device write; a rejection keeps the active session. A
device-transfer failure after that point ends the engine, which the server restarts.
- The model fingerprint samples every model file the engine loads, by role (GGUF shards and overrides, PLE, the
pack's own files, every expert file as the loader resolved it per layer and role, the MTP files); the config fingerprint covers the engine version, backend, KV settings, context,
the resolved RoPE configuration, a digest of the loaded control vectors and the arithmetic switches.
- The hashes detect accidental corruption; they do not authenticate files. Restore only files this engine wrote.
## Limits
- One slot (slot 0), one GPU: refused with `--layer-split` and with `--peer-device`.
- Requires `--prompt-cache` > 0; the separate RAM parked-conversation cache need not be on.
- A restore does not park the outgoing session; only the deepest checkpoint is saved.
- Session files are kept until deleted (about 1.2 GB per 63K-token conversation).
- All requests keep the Host and API-key checks; slot actions also require JSON and no `Origin`, the server's own or
a trusted one, also when an API key is set.
## Tested on
- Linux + NVIDIA: measured (above).
- Windows: the file I/O has a Win32 branch (`CreateFileW` with `CREATE_NEW`, no-follow open, `FlushFileBuffers`,
`MoveFileExW`). Its CPU file-I/O unit test passed as a 32-bit Windows executable built with zig 0.16 (clang,
`x86-windows-gnu`) under Wine 10.0; this is not a Windows engine or GPU validation. Not built with MSVC, not run on
real Windows, Unicode model paths and large files not tested there.
- AMD (HIP): no new GPU code; device reads go through the existing helpers that `hip_compat` maps. Not built
or run with `STRATA_ENABLE_HIP=ON`.
If you are on Windows or AMD and would like to try it, a run of the unit test and one save/restore of a long chat
would help a lot. I have a branch ready on top of v0.1.38 and will open a PR shortly; two robustness items are still being finished
there (the RAM preflight of a save runs before its first copy, and progress lines cover small saves and the
flush phases).
Sur le site
Liens install, modèles, releases.