Pull requests / #668

#668 Session files: save and restore a conversation to disk (POST /slots/0?action=save|restore)

closed · @maverde73 · 0 评论 · 在 GitHub 查看

Setup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentationWindowsLinux

描述

Saves the conversation the engine holds to one file and restores it later, also after a restart of the same
engine version with the same model and settings. At 63K tokens on an RTX 4070 Ti the restore took 1.05 s once the
model was loaded (engine startup itself takes about 70-95 s) instead of re-reading the prompt for 25.5 s, and the next
turn gave the same 32 token IDs as without a restart. Files are bound to the engine version, model inputs and settings, must be trusted (the hashes
detect corruption, they do not authenticate), and there is one slot, slot 0.

Proposed in #661. Base: v0.1.39 (6f32ec0); rebased from v0.1.38 (see "Rebased on v0.1.39" below). Opt-in: without
`--slot-save-path` nothing changes - in three alternating pairs at 3.4K, 29.6K and 63K there was no measurable
prefill, decode or peak-VRAM difference and the output token IDs were identical in all nine pairs (measured on the
v0.1.38 base; the session code adds nothing to the inference path).

## What it does

- Engine: `strata --serve` takes `SAVE <path>` and `RESTORE <path>` between requests and answers
  `SAVED <tokens> <bytes> <ms>`, `RESTORED <tokens> <bytes> <ms>`, `SERR <kind> <published 0|1> <reason>` (failed;
  the engine and its session as they were; kind = `invalid`, `storage`, `memory` or `io`) or `FATAL <reason>` (a
  restore transfer failed after the device writes began; the engine exits). `SESSION <done> <total>` follows every
  block of the file that moved (at most 16 MiB; the last partial block and a small file included). `SWAIT <phase>
  <seconds>` comes before a step that blocks in one call (fingerprint, state capture, file flush, rename + folder
  flush, validation, device transfer): the engine's watchdog and the server allow that step those seconds (60 + 1 per
  4 MiB, at most 3600), then it is stuck; the next line clears the allowance. The server checks every line and ends an
  engine that answers out of protocol (malformed, negative or non-finite numbers, progress going back, an allowance
  outside 1..3600). Server and engine change together: an older server does not know `SERR`/`SWAIT`.
- Server: llama-server-compatible save and restore requests and response fields for slot 0, opt-in with
  `--slot-save-path DIR` (or `"slot_save_path"` in the config), resolved once to an absolute path and created with
  mode 0700. NAME must be one plain file name, portable to Windows (no drive, stream, device name, trailing dot or
  space). `501` disabled, `400` refused name/slot/action or an invalid file (corrupt, another model or
  configuration, over this session's limits), `404` missing file, `503` model not loaded, RAM preflight refused or
  allocation failed, `507` no disk space (preflight, or the OS reports no space/quota), `500` other I/O failures and
  an engine that ended, was silent or answered out of protocol (it is restarted on the next request). The status
  comes from the engine's category (`error.kind`), never from the message text. Slot actions
  queue behind running requests, show in `/status` and count as activity for the idle unload.
- All requests keep the Host and API-key checks. Slot actions additionally require JSON and no `Origin`, the
  server's own or an explicitly trusted one; this Origin rule also applies when an API key is configured.
- Save: an exclusively created temporary file with a unique hidden name beside the target (file mode 0600 on
  POSIX), flushed (`fsync`/`FlushFileBuffers`), renamed (`rename`/`MoveFileExW`), folder flushed on POSIX. A save
  that fails before the rename keeps the old file and removes only its own temporary file. If the folder flush fails
  after the rename, the save fails with `error.published: true`: the new file has replaced the old one, but its name
  may not survive a power loss. A folder that cannot be flushed (`EINVAL`) is logged, not an error. Before writing,
  the save needs room for the whole new file plus `--session-min-free-mib` (engine argument, default 4096, 0 = off;
  5.12 GiB free for a 63K file), also when replacing a file; and its RAM preflight - the checkpoint copies, the copied live state, the source
  directory and the write buffer, above the parking floor - runs before anything is copied. Both are preflights, not
  reservations; RAM is the host's available memory, not a cgroup limit.
- Restore: no-follow open (`O_NOFOLLOW`, `FILE_FLAG_OPEN_REPARSE_POINT`), only a regular file with one link. File
  and runtime validation (size against the largest this session can restore, header, model and config
  fingerprints before the payload is parsed, a RAM preflight for the parse's peak, geometry and layer range, every count against the bytes left
  and this session's exact limits - tokens, checkpoints, layers, each state array and K/V part - before allocating
  it, payload hash, snapshot validation) complete
  before any device write; a rejection keeps the active session. A device-transfer failure is fatal.
- Identity: the model fingerprint samples every model file the engine loads, by role (GGUF shards, overrides,
  PLE, the pack's own files, the MTP files) and every file the expert source actually read its experts from, per
  layer and role as `native_experts.txt` resolved it; each (role, file) pair counts, each file is read once; the config fingerprint holds typed, length-prefixed fields: engine
  version, backend, KV settings, context, the resolved RoPE configuration, a digest of the uploaded control-vector
  tables (content, exact scale, mode, layers, direction) and the arithmetic switches. The version is the release
  string: another version is refused, two builds of one version are not told apart.
- K/V is streamed from where it lives (device or pinned host pool) in 16 MiB aligned blocks, no full host copy.
- Format v1, little-endian, documented in docs/DETAILS.md, with a golden test of the fixed sample's bytes.

Refused: `--layer-split`, `--peer-device`, `--batch` / `"parallel"` (#465; the server answers `501`), `--prompt-cache 0`.

<details><summary>Files</summary>

| File | Change |
| --- | --- |
| `include/strata/core/conversation_file.hpp`, `src/core/conversation_file.cpp` | new: format, I/O (POSIX and Win32), identity, fail-closed read |
| `src/core/conversation_file_test.cpp` | new: 185 CPU checks (182 under ASan, which skips the address-space-limit case) |
| `include/strata/core/conversation_snapshot.hpp`, `src/core/conversation_snapshot.cpp`, `src/core/conversation_state.cpp` | K/V layers as streamed sources |
| `src/core/conversation_snapshot_test.cpp` | streamed file equals the captured-snapshot file (GPU) |
| `src/program/generate.cpp`, `include/strata/core/progress.hpp` | `SAVE`/`RESTORE`, fingerprints, `--session-min-free-mib`, the watchdog's bounded allowance for blocking steps |
| `include/strata/core/expert_source.hpp`, `src/core/expert_source.cpp` | the files the expert source read, for the fingerprint |
| `serve/server.py`, `serve/test_slots.py`, `serve/test_security.py` | the slot API, name policy, protocol checks, typed errors, `SWAIT`, status; 30 + security tests |
| `CMakeLists.txt`, `docs/DETAILS.md` | build, docs |

</details>

## How to test

```bash
# CPU only
cmake -B build -DSTRATA_BUILD_CONVERSATION_TESTS=ON ... && cmake --build build --target conversation_file_test
./build/conversation_file_test
STRATA_SESSION_BUFFERED=1 ./build/conversation_file_test    # buffered path
python -m unittest serve.test_slots serve.test_security      # mock engine
# GPU
./build/conversation_snapshot_test
# end to end
python serve/server.py --config strata.json --slot-save-path /path/to/sessions
# a long chat, then save; restart the server; restore; then send the same chat's next message
curl -X POST "http://127.0.0.1:8080/slots/0?action=save" -H "Content-Type: application/json" -d '{"filename":"a.bin"}'
```

A request whose messages continue the restored conversation reuses the restored state or its checkpoint; a short
tail may be read again. Clients still send their messages and images.

## Rebased on v0.1.39

Two textual conflicts, both resolved keeping upstream's code: the `/v1/responses` dispatch in `serve/server.py` (the
slot route is now one more `elif` after it) and the `--conversation-cache-min-free-mib` help line in `generate.cpp`
(next to `--vram-elastic` / `--batch`); `ArenaExpertSource::close()` merged automatically except for one hunk (the new
`STRATA_ARENA_MMAP` unmap stays, the PR's `inputs_.clear()` added after it). One commit on top for two things
v0.1.39 changed underneath the PR:

- #465 batch slots: parallel requests do not take the server's FIFO, so a slot save/restore could interleave with a
  request being admitted. With `--batch` the engine refuses `SAVE`/`RESTORE` (`SERR invalid`) and the server answers
  `501` before sending anything (new test in `serve/test_slots.py`; docs updated).
- `STRATA_ARENA_MMAP=1`: the new mapped-arena path returned before recording the experts file it read, so the model
  fingerprint would have missed `experts.bin`; it now records it like the pinned arena.

#583 (the byte-budget ring, on by default) only sizes the prompt path's streamed ring; nothing it changes is in the
saved state, so the file format (v1) and its golden bytes are unchanged.

Checked on v0.1.39 (`6f32ec0` + this branch), the same machines as below; unmodified v0.1.39 built and run the same way
for comparison:

- CUDA (RTX 4070 Ti machine, gcc-15, sm_89): full build, 0 new warnings (47 against 48: the unused `argmax` removed).
  `conversation_file_test` 185, `conversation_cache_test` 4191, `conversation_memory_test` 23,
  `conversation_validation_test` 972 (host-only), `file_expert_source_test`, `expert_profile_save_test`,
  `pinned_shared_test` pass. The GPU there was in use, so `conversation_snapshot_test` ran on AMD (below).
- Server suite (`serve/test_*.py`): 299 tests run on the branch, 268 on unmodified v0.1.39 (+30 `test_slots`, +1
  `test_security`); in both the same one failure, not this PR's: `test_responses.OverHttp.test_json_schema_text_format`.
- HIP (RX 6650 XT, as below): 0 new warnings (1196 against 1197); `conversation_snapshot_test` 3,901 checks on the GPU,
  6 runs of 6 (unmodified v0.1.39: 2,169); `conversation_file_test` 185 (also buffered); `ctest -R conversation_` 5/5.
- Not repeated on v0.1.39: the end-to-end 63K save/restore with the model (numbers below are from the v0.1.38 base).

## Tested on (v0.1.38 base)

Linux + NVIDIA: RTX 4070 Ti 12 GB, Ryzen 9 5900X, 64 GB RAM, NVMe ext4, CUDA 13.1, gcc-15, sm_89, IQ3_XXS pack, the
setup's config (spec 4 + MTP, `--kv int8`), binary sha256 `3bbe4fc38ab3...` (this HEAD's code; the last commit
changes docs only):

- builds; the one warning in the build log (`verify.cpp`, unused variable) is in a file this change does not touch;
  `conversation_file_test` 185/185 (default and `STRATA_SESSION_BUFFERED=1`), `conversation_cache/memory/validation_test`
  pass; `conversation_snapshot_test` 3,621 checks pass on the GPU. On a second Linux PC (gcc 15) the file test also
  builds with `-Wall -Wextra -Werror` and passes under ASan+UBSan;
- server suite: 227 run, 9 skipped, no failures (vanilla v0.1.38: 197 run, 9 skipped); the tools suite fails the
  same way with and without the branch in this Linux environment (it needs llama.cpp's gguf-py), so it is not
  claimed as passing;
- 12,287 teacher-forced logprob rows from a 12,288-token input through the short-read path (`--short-read 13000`)
  are byte-identical to v0.1.38's; this does not establish parity of every batched-prefill or restore path.

<details><summary>End to end at 63K</summary>

62,993-token prompt + 32 tokens, `--max-context 65536`. Single runs:

| | |
| --- | --- |
| file | 1,198,691,396 bytes, mode 0600 |
| save (flushed) | 653.9 ms; 73 `SESSION` lines, at most 16 MiB apart; `SWAIT` allowances: capture 88 s, flush 346 s, publish 346 s |
| 10 more saves (1,199,613,632-byte state after a further turn) | replacing one file 594.0-611.4 ms; to new names 797.1-1181.5 ms |
|

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。