Pull requests / #41
#41 serve: preserve primary conversation state across auxiliary calls
closed · draft · @midhatn · 0 comments · View on GitHub
Server & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Description
An agent using one local model for both its primary conversation and auxiliary approval/summary calls can lose its primary KV state between turns. The ordinary conversation checkpoints do not preserve an independent positional KV cache, so returning from an unrelated side call can reread thousands of tokens. This adds an opt-in, bounded host-RAM snapshot around explicitly marked auxiliary requests. A primary → auxiliary → primary sequence restores the primary only when the token/image prefix and control-vector setting match; rewritten history falls back to normal prefill. The feature defaults off and does not add parallel model instances. ### Changes - `--conversation-cache-mib N` parks one complete mutable primary state, including QSA device/host KV, GDN recurrence/convolution, PLE history, MTP state and the associated checkpoints/token/image metadata. Model weights and expert placement stay in place; device addresses remain stable. - The API adapter forwards a literal `strata_auxiliary: true` as `aux=1` for text and image requests. Further marked auxiliaries do not displace the parked primary, even if their prompts are longer. Reasoning level and prompt text are not used to infer auxiliary intent. - Admission enforces the byte cap and a 2,560 MiB available-physical-RAM floor immediately after allocation. Allocation/copy failure falls back to cold prefill; partial restore failure terminates the engine rather than continuing with inconsistent state. - The adapter prefers the running engine's INFO version to an adjacent `BUILD.json`. This fixes a custom 0.1.13 executable being displayed as 0.1.12 when it shares a directory with a retained official fallback. Older engines still use the manifest fallback. - Add API/protocol tests, a standalone real-CUDA/host round-trip test with checks active in Release, and usage/validation documentation in `docs/AUXILIARY_STATE_CACHE.md`. ### Validation Based on upstream `b89c989a7155e984e544ddd90d1038dda7da9e3d`, including the prompt-cache correctness fixes after 0.1.13. The existing patch was also compatibility-tested against v0.1.16; the published implementation has not yet been rebased onto that release. Hardware used for bounded validation: Ryzen 9 7940HS, RTX 4070 Laptop 8 GB, and 64 GB RAM, running Windows. - Server/adapter tests and the release engine build passed. - The native host/device snapshot test passed round-trip restoration, replacement, admission bounds and invalidation checks. - Bounded checks covered same-image restoration, changed-image invalidation, an auxiliary longer than its primary, rewritten-history invalidation, and subsequent matching-prefix restoration. - Code outputs were independently executed and checked; tool-call, reasoning and auxiliary-request routing checks also passed. - A long-context known-answer continuation after marked side calls crossed the resident KV window and preserved the expected answer while reusing the primary prefix. These are bounded checks, not the requested repeated streamed-KV soak or proof of arbitrary agent-task success. They do not establish bitwise logit parity, cross-platform correctness or a general generation-speed improvement. ### Scope for review This is a **draft** for review of the API and state ownership. Physical-memory admission currently uses Windows telemetry; other platforms decline admission safely. Linux/cgroup-aware admission is not implemented or validated. The snapshot cap excludes existing small checkpoints, and its RAM-floor check is not a continuous system-wide reservation. Bitwise logit parity across all model/configuration combinations is not claimed. The cache protects one primary around explicitly marked side calls. It is not persistent storage, arbitrary multi-session caching or a solution for rewritten compression prefixes. The auxiliary inference itself still runs. Deployment-specific configuration and agent delegation changes are outside this PR. Related to the prompt-cache work in #39, but this snapshots the complete mutable state around explicit auxiliary calls rather than changing checkpoint retention or telemetry. No maintainer action or merging is assumed.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.