Pull requests / #175

#175 Preserve independent conversation caches across interleaved requests

closed · @b7216309-jpg · 0 评论 · 在 GitHub 查看

BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindows

描述

## Motivation



Alternating a long conversation with a short background request (for example, memory extraction) overwrites its positional KV branch. Prefix checkpoints can retain running state, but cannot resume cells overwritten by another conversation. Returning to the original chat therefore rereads its large prompt.



This adds explicit independent cache slots for single-GPU sessions, while keeping one execution lane and one resident model/session arena.



## Changes



- Accept `strata_cache_slot` (integer 0–3, default 0) on both chat-completions and messages APIs. Invalid selections fail before SSE headers are sent.

- Snapshot inactive main/MTP positional KV and pooled index data to delete-on-close temporary files. Preserve live token/image identity, recurrent/PLE/index-tail state, prefix checkpoints and the control-vector setting.

- Restore into the existing arena, reset streamed GPU page mappings, and retain upstream's pinned-root/LRU checkpoint policy within each slot.

- Recover the latest valid checkpoint when rotating away from a cancelled request.

- Advertise `cache_slots` / `cache_storage` in engine INFO. Layer-split multi-GPU engines advertise one slot and reject nonzero selections instead of taking an incomplete snapshot.

- Add API regression tests, a standalone live-model correctness script and documentation of the protocol/resource costs.



Ported onto current main (`a790805`, 0.1.27). Existing upstream cached-token reporting and image sampling behavior are preserved; this does not duplicate those newer fixes.



## Validation



Built the CUDA `strata` target on Windows using MSVC, CUDA 13.3 and `CMAKE_CUDA_ARCHITECTURES=89`.



`python -m unittest discover -s serve`: **70 tests passed, no skips**, with `STRATA_TOKENIZER` pointing to the installed pack tokenizer.



Ran the included `python tools/check_cache_slots.py --lines 1700` against this PR build on an RTX 4070 Ti / Ryzen 7700X / 64 GB RAM, using Qwen3.8-Flash-Next IQ2_XS, int8 KV, a 262,144-token context, 32,768 resident KV cells, MTP enabled and `--adapt-swaps 0`:



| Request | Prompt tokens | Cached tokens | Wall time |

| --- | ---: | ---: | ---: |

| A cold | 35,771 | 0 | 19.985 s |

| B cold | 35,771 | 0 | 19.578 s |

| A restored | 35,771 | 35,764 | 0.843 s |

| B restored | 35,771 | 35,764 | 0.782 s |

| A independent cold | 35,771 | 0 | 19.578 s |



Cold/restored greedy answers matched exactly and retrieved both requested distant record markers. After disconnecting a streaming generation, the restored slot also matched an independent cold run (57 cached tokens versus 0). The script asserts reuse from reported token counts, not elapsed time. These are synthetic local measurements, not an end-to-end application benchmark.



## Scope and tradeoffs



- Opt-in slot selection; omission retains slot 0 and avoids snapshot I/O when requests remain there. This does not enable parallel generation or implement automatic conversation identification/eviction.

- Up to three inactive snapshot files hold positional KV, recurrent state and prefix checkpoint payloads. The active slot retains its usual RAM state; small inactive descriptors remain resident. Weights and GPU arenas are not duplicated.

- Server-process lifetime only: restarting clears the caches. Disk/CUDA snapshot failures surface as engine errors rather than using partial state.

- Slot numbers are server-global, not per-client namespaces. Clients must coordinate assignments and continue sending their full message history.

- Live-tested on single-GPU NVIDIA/int8 KV with text requests. Multi-GPU rotation is explicitly unsupported; HIP, other KV formats and image-cache rotation have not been live-validated.





### SSD checkpoint offloading update



Inactive recurrent state, token/image identity and prefix checkpoints are now appended to the positional snapshot file and their vector allocations are released. Active state remains in RAM. Small descriptors remain resident, and filesystem cache is reclaimable by the OS. In a Windows/NVMe test, peak additional private-memory growth fell from 899.9 MiB to 34.2 MiB; restored-slot switches took 437–938 ms. The 35k-token retrieval/rotation/cancellation parity check and the Little Bot streaming relay passed. Use a separate model server for diagnostics: cache slot IDs are server-global.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。