Pull requests / #861
#861 serve: "strata_checkpoint": false lets a one-shot request skip its conversation checkpoint (#830)
closed · @brenoperucchi · 0 comentarios · En GitHub
Setup & installServer & APINVIDIA / CUDAModels & quantsSecurityDocumentationWindows
Descripción
For #830. A request can now send `"strata_checkpoint": false` when no later request will continue it, such as a classification call or a probe. For that request only, the engine: - saves no checkpoint at the last turn boundary, so the prompt after the reused prefix (or the root) is read in one part instead of two; - saves no periodic checkpoint every `--prompt-cache-every` tokens; - leaves no live session to continue, so `--conversation-cache` doesn't park it. It still resumes from a checkpoint it matches, and it still saves the system-prompt root when that reaches `--prompt-cache-root`. Without the field, or with `true`, nothing changes. The server passes it to the engine as `ckpt=0`, next to `cvec=0`; an engine without the key skips it. I didn't use `cache_prompt: false` for this. In llama.cpp it means "don't reuse the cache", and these calls need to keep reusing the root. ## Changes - `src/program/generate.cpp`: parse `ckpt=0`. With it, `turn_at` is dropped (the root is still looked for up to the last turn), `pp_next_check` stays at `INT64_MAX`, and `live_ok` stays false. `STRATA_STATE_HASH` still prints for the request. - `serve/server.py`: `"strata_checkpoint": false` adds ` ckpt=0` to the request keys. - `serve/test_server.py`: the key is sent only for `false`. `cache_prompt: false` doesn't send it. - `docs/DETAILS.md`: one paragraph under the conversation cache. ## Measurements I built this branch myself on Windows (CUDA 13.0, sm_120, Release) and ran both arms on the same binary and the same server start, alternating A B A B. Setup: RTX 5090 32 GB, Ryzen 9 5950X, 96 GB DDR4-3200, Swift IQ3_XXS, `--prefill auto:32768 --max-context 32768 --kv int8 --pcie-frac 0.20 --spec 4 --spec-min-p 0.70 --prompt-cache-root 64`, `STRATA_IQ_MT_MIN=1`. The calls have the shape from #830: a fixed system prompt of about 70 tokens and 150 to 450 tokens of public code (excerpts of v0.1.32's `generate.cpp`), `max_tokens` 1, thinking off, 12 calls per size with the first 2 skipped. The client ran on the same PC. Medians of the second pair: | User message | Tokens read | `prompt_ms` without | `prompt_ms` with `false` | Wall without | Wall with `false` | |---|---|---|---|---|---| | ~150 tokens | 160 | 335 | 293 | 350 | 304 | | ~300 tokens | 310 | 400 | 362 | 412 | 379 | | ~450 tokens | 418 | 427 | 388 | 440 | 405 | The table shows the smallest of the three pairs. The `prompt_ms` saving per pair was 46/47/46 ms (A1/B1), 42/38/39 ms (A2/B2) and 47/44/45 ms (traced pair). Both arms reused the 69-token root on every call after the first one of each start, which builds it. All 216 one-token answers were `SAFE` in both arms, so they only show that the label didn't flip. In a separate start with `STRATA_TRACE=1`, after the first call, the default arm reads each prompt in two parts (`(batched)`, then 6 tokens `(windows)`) and ends with `2 checkpoints` kept. With `false` there is one `(batched)` read, no `(windows)` read, and `1 checkpoints` (the root). ## State parity To check the state itself, I ran two more starts with `--adapt-swaps 0` and `STRATA_STATE_HASH=1`, the second one also with `STRATA_CKPT_REREAD=1`, and compared the hash printed after each request. For the 18 requests with `"strata_checkpoint": false`, `L`, `gdn`, `ple`, `tail`, `pooled`, `kv` and `mtp` were identical between restoring the root and reading it again (only `stale` differed, in 2 of them). The 18 default requests differed in every field, which I'd expect: `STRATA_CKPT_REREAD` reads the 6-token header batched instead of through the windows. With `false` the header goes batched in both starts, so that pair compares the same reads. ## Not covered - I only ran it on one GPU. I didn't test a layer split, `--conversation-cache-mib` or #718's disk parking; for parking, the change is only that `live_ok` stays false, and #718's park checks `live_ok` too. - `--batch` / `"parallel"`: a `false` request that continues in a slot still leaves that slot cached. It does no extra copy and the next probe doesn't match that slot, but this change doesn't cover it. - The SYCL port (`sycl/src/program/generate.cpp`) has the same logic and isn't changed here. It ignores `ckpt=0` until it is ported. - If #614's near-tail checkpoint goes in, a `false` request should skip that one too. - With `false` the 6-token assistant header goes through the batched path instead of the windows, so its rounding differs slightly from the default arm, as it does with `--prompt-cache 0`. A parity test should compare the same arm against itself, as above. I can rerun anything you want on this machine, for example `tools/conversation_cache_parity.py` with some `false` requests mixed in.
En el sitio
Enlaces a install, modelos, releases.