Pull requests / #189
#189 Preserve alternating conversations with a shared snapshot core and bounded RAM cache
closed · @jeremiahritchey · 0 コメント · GitHub で見る
BenchmarksSetup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsSecurityWindowsLinux
本文
Alternating agent conversations currently lose reusable state when another conversation takes over the session. This adds an opt-in, bounded RAM cache that selects the longest compatible prefix and restores it through one shared snapshot implementation. Requests remain serial, require no client session markers, and reuse existing GPU allocations.
Refs Niko1221/Strata#57; related to Niko1221/Strata#143. This is the shared-core/RAM part of the agreed split with #41 and #52, based on current main, 0.1.27 (`a790805`). The optional disk tier is reviewed separately in [upstream PR #190](https://github.com/Niko1221/Strata/pull/190), which depends on this core landing first. Its description links the NVMe-only comparison. General Pi benchmark tooling is excluded from this PR's final diff.
The snapshot covers used main and draft KV, GDN/PLE, indexer state and retained checkpoints. RAM admission checks both capacity and physical headroom on Linux/Windows; invalid snapshots fall back before application. A GPU transfer or synchronization failure stops the engine rather than continuing from partially restored state. K8V4 is represented explicitly, and opt-in draft read-back verifies authoritative and resident-ring bytes.
Validation on Linux, EPYC 7532 / RTX 4090, GCC 15.2 / CUDA 13.4, IQ3_S:
- Component CTests and 1,781 real-GPU snapshot checks pass; transfer diagnostics pass 1,452 ASan/UBSan host checks.
- Paired INT8 streamed/ring and INT8/K8V4 batched-prefill gates pass known-answer, exact-token and main-state comparisons, with draft read-back evidence.
- Pressure, oversized-entry, exchange, admission-denial, synthetic image/grid and add/project steering isolation gates pass.
- A 30-cycle soak reaches 119,987 prompt tokens; twelve HTTP requests cover reuse, eviction, disconnect and recovery.
- The cache harness suite passes 34 tests under normal and optimized Python. Earlier frontend validation passed 66 tests with three skipped.
Windows code was reviewed but not locally executed; HIP, real vision encoding and the full Coder model remain untested. Whole-conversation multi-GPU parking and hybrid K8V4 streaming/rings are unsupported. These correctness results do not establish an overall Pi speedup or production-context performance.
@midhatn's Windows admission test retains its original authorship in `b319d43`. Thanks to @maedoc for the NVMe collaboration and @QilinWan for the identity/coordinate review.
The validation summary above applies to the tested 0.1.27 engine. Detailed design/validation notes and raw model/build logs are retained locally; planning documents are excluded from this PR.
Additional default-off regression (2026-09-30): untouched upstream `a790805`, this core build, and the disk-enabled #190 build pass 72 requests total (eight per engine per mode): INT8 resident ring/spec-4 with a 43,170-token prompt, batched INT8, and batched K8V4. All conversation-cache flags are omitted from every arm. Identical assets, effective settings, and request tokens produce exact output-token, comparable main-state, prefix-reuse and known-answer parity. The new verifier passes six tests including negative controls under normal and optimized Python; the combined cache harness suite passes 34 tests. These are correctness checks, not speed measurements.
Reviewer follow-up: the pressure gate reports each snapshot size and a usable budget range. Restored conversations retain uniquely owned K/V buffers; subsequent parks copy changed pages while capturing running state and checkpoints afresh. Retained capacity counts against the RAM budget, and full capture remains the fallback when growth would evict another conversation. See the [independent validation and suggestions](https://github.com/Niko1221/Strata/pull/189#issuecomment-5903850856). The paired 12-request growth/rewind gate passes batched INT8, batched K8V4 and INT8 resident-ring/spec-4 at 61,259 prompt tokens. On this machine, repeated long-conversation parking takes 0.37–0.52 s versus 0.97–1.12 s with full capture, avoiding 4.3 GiB of KV copying across the sequence. Both arms use the same executable and enabled cache, with exact output/main-state and restored draft read-back checks. This measures parking, not end-to-end agent performance. A separate 835 MiB run exercises full-capture fallback for the largest growth step while smaller parks retain pages; both arms finish with zero evictions and identical output/state.
Follow-up validation at `f1b0c16` (core) / `18c475e` (NVMe): all 23 final GPU/model jobs pass, including growth/rewind, tight-budget fallback, cache-on/off parity, isolation, HTTP recovery, the 30-cycle soak, three disk lifecycles, and the 72-request upstream/default-off matrix. Core and NVMe each pass 1,781 GPU snapshot checks. Host sanitizer checks include 4,149 buffer/cache policy checks, 1,452 transfer checks, and 6,732 disk-codec checks. Windows execution remains untested locally.
Portability follow-up (`81cb973`): cache harness text reads/writes explicitly use UTF-8, and isolation fixtures use absolute paths so an engine configured with a different working directory can locate them. This addresses the Windows tester’s two inline findings. All 34 cache harness tests pass under normal and optimized Python. CPU-only reproductions fail on the prior code and pass after the fix for simulated cp1252 defaults and differing engine working directories (image/add/project). These checks do not constitute a Windows engine rerun.
<details>
<summary>Reproduction commands (Linux, IQ3_S / RTX 4090)</summary>
Run from this PR branch's root, with the model installation's Python environment activated. These commands reproduce the Linux RTX 4090 settings (CUDA 13.4, SM89, IQ3_S); replace the two asset paths. The model config must contain absolute asset paths. Model runs and GPU fixtures require an exclusive GPU window. Each model output directory must be new. This is correctness validation, not a throughput benchmark.
```sh
export STRATA_CONFIG=/absolute/path/to/strata-iq3_s.json
export STRATA_LLAMA=/absolute/path/to/llama.cpp
export STRATA_GGUF_PY="$STRATA_LLAMA/gguf-py"
export CUDACXX=/usr/local/cuda-13.4/bin/nvcc
```
The dependency is the setup-pinned llama.cpp commit `3cf03257f219afbe7334045ff7c6a06ac68c627d`. Preserve your config's model/tokenizer/profile paths while selecting the tested context, residency and expert-cache settings:
```sh
python - <<'PY'
import copy
import json
import os
from pathlib import Path
source = json.loads(Path(os.environ['STRATA_CONFIG']).read_text())
keys = {'--max-context', '--kv', '--kv-resident', '--expert-cache', '--pcie-frac'}
args, i = [], 0
while i < len(source['args']):
flag = source['args'][i]
if flag in keys or flag.startswith('--conversation-cache-'):
i += 2
else:
args.append(flag)
i += 1
out = Path('logs/pr-repro')
out.mkdir(parents=True, exist_ok=True)
for name, context, kv, resident in [('int8-ring', 131072, 'int8', 32768),
('int8-batched', 16384, 'int8', 0),
('k8v4-batched', 16384, 'k8v4', 0)]:
cfg = copy.deepcopy(source)
cfg['args'] = args + ['--max-context', str(context), '--kv', kv,
'--kv-resident', str(resident), '--expert-cache', '6000',
'--pcie-frac', '0.55']
cfg['cwd'] = str(Path.cwd())
cfg.pop('log', None)
(out / (name + '.json')).write_text(json.dumps(cfg, indent=2) + '\n')
PY
```
Build and host checks:
```sh
cmake -S . -B build-review-cuda -DCMAKE_BUILD_TYPE=Release \
-DSTRATA_ENABLE_CUDA=ON -DSTRATA_ENABLE_HIP=OFF -DCMAKE_CUDA_ARCHITECTURES=89 \
-DSTRATA_PORTABLE=ON -DSTRATA_NATIVE_EXPERTS=ON -DSTRATA_PREFILL_MMQ=OFF \
-DSTRATA_BUILD_TESTS=OFF -DSTRATA_BUILD_CONVERSATION_TESTS=ON \
-DSTRATA_GGML_DIR="$STRATA_LLAMA"
cmake --build build-review-cuda -j 2 --target strata conversation_cache_test \
conversation_memory_test conversation_snapshot_test conversation_validation_test \
conversation_transfer_test conv_cache_test sampler_parity
CUDA_VISIBLE_DEVICES=-1 ctest --test-dir build-review-cuda \
-R '^(conversation_cache_test|conversation_memory_test|conversation_validation_test|conversation_transfer_test|conv_cache_test)$' --output-on-failure
python -m unittest tools.test_conversation_cache_parity tools.test_conversation_cache_isolation \
tools.test_conversation_cache_soak tools.test_conversation_cache_http_smoke tools.test_conversation_cache_disabled \
tools.test_conversation_cache_growth
python -O -m unittest discover -s tools -p 'test_conversation_cache_*.py'
```
GPU/model checks, only in the exclusive window:
```sh
ctest --test-dir build-review-cuda -R '^(conversation_snapshot_test|sampler_parity)$' --output-on-failure
export STRATA_SNAPSHOT_VERIFY=1
export STRATA_MTP_BATCH=1
numactl --interleave=all --physcpubind=0-31 python tools/conversation_cache_parity.py \
--config logs/pr-repro/int8-ring.json --engine build-review-cuda/strata \
--output logs/pr-repro/ram-ring --run
numactl --interleave=all --physcpubind=0-31 python tools/conversation_cache_parity.py \
--config logs/pr-repro/int8-ring.json --engine build-review-cuda/strata \
--spec 4 --paragraphs 3600 --output logs/pr-repro/ram-spec4 --run
for mode in int8-batched k8v4-batched; do
numactl --interleave=all --physcpubind=0-31 python tools/conversation_cache_parity.py \
--config "logs/pr-repro/$mode.json" --engine build-review-cuda/strata \
--output "logs/pr-repro/ram-$mode" --run
done
for scenario in pressure exchange; do
numactl --interleave=all --physcpubind=0-31 python tools/conversation_cache_parity.py \
--config logs/pr-repro/int8-ring.json --engine build-review-cuda/strata \
--scenario "$scenario" --cache-mib 400 --output "logs/pr-repro/$scenario" --run
done
numactl --interleave=all --physcpubind=0-31 python tools/conversation_cache_parity.py \
--config logs/pr-repro/int8-ring.json --engine build-review-cuda/strata \
--scenario oversized --cache-mib 1 --output logs/pr-repro/oversized --run
numactl --interleave=all --physcpubind=0-31 python tools/conversation_cache_parity.py \
--config logs/pr-repro/int8-ring.json --engine build-review-cuda/strata \
--scenario admission --min-free-mib 1048576 --output logs/pr-repro/admission --run
for scenario in image add project; do
numactl --interleave=all --physcpubind=0-31 python tools/conversation_cache_isolation.py \
--config logs/pr-repro/int8-ring.json --engine build-review-cuda/strata \
--scenario "$scenario" --output "logs/pr-repro/isolation-$scenario" --run
done
numactl --interleave=all --physcpubind=0-31 python tools/conversation_cache_soak.py \
--config logs/pr-repro/int8-ring.json --engine build-review-cuda/strata \
--output logs/pr-repro/soak --run
```
For the default-off regression, build the untouched reference with matching compiler/backend flags. Use a new sibling path:
```sh
git worktree add --detach ../Strata-upstream-reference a79080535d1b2a71a3419a0d97d8e7dca194b0f1
cmake -S ../Strata-upstream-reference -B ../Strata-upstream-reference/build-review-cuda \
-DCMAKE_BUILD_TYPE=Release -DSTRATA_ENABLE_CUDA=ON -DSTRATA_ENABLE_HIP=OFF \
-DCMAKE_CUDA_ARCHITECTURES=89 -DSTRATA_PORTABLE=ON -DSTRATA_NATIVE_EXPERTS=ON \
-DSTRATA_PREFILL_MMQ=OFF -DSTRATA_BUILD_TESTS=OFF -DSTRATA_GGML_DIR="$STRATA_LLAMA"
cmake --build ../Strata-upstream-reference/build-review-cuda --target strata -j 2
numactl --interleave=all --physcpubind=0-31 python tools/conversation_cache_disabled.py \
--config logs/pr-repro/int8-ring.json \
--upstream ../Strata-upstream-reference/build-review-cuda/strata \
--candidate build-review-cuda/strata --paragraphs 2600 --spec 4 \
--output logs/pr-repro/disabled-int8-ring --run
for mode in int8-batched k8v4-batched; do
numactl --interleave=all --physcpubind=0-31 python tools/conversation_cache_disabled.py \
--config "logs/pr-repro/$mode.json" \
--upstream ../Strata-upstream-関連リンク
インストール・モデル・リリースへの站内リンク。