Pull requests / #275
#275 Add bounded disk persistence for evicted conversations
closed · @jeremiahritchey · 0 comentários · No GitHub
BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Descrição
RAM-evicted conversations can retain reusable state on disk and return after an engine restart. This adds an opt-in disk tier with byte/entry quotas, using the snapshot representation and GPU restore path released in #189.
Replaces #190 following the [request to rebase onto the shipped core](https://github.com/Niko1221/Strata/pull/190#issuecomment-5917508859). Based on v0.1.30 (`30ec18e`), including `e46ccca`'s session carve. The diff contains only the dependent disk changes; #189 has shipped. Refs #57 and #52. Delta writes remain a possible follow-up from #265.
- Spill only on RAM eviction, with bounded staging, an exclusive process lock, interrupted-file cleanup and atomic publication. A continuing or latest active turn is not automatically durable.
- Select compatible prefixes across active state, RAM and disk. Validate the full envelope and integrity before GPU application; rejected/corrupt/unadmitted entries fall back to prompt processing.
- The v2 envelope preserves the session layer range introduced by #216. Named identity fingerprints bind full asset contents, engine, geometry, per-layer expert quantization, KV layout/residency, resolved RoPE, split/device/runtime and inference settings. Log the first differing field before reading state.
- Disabled persistence performs no cache-directory I/O. OpenSSL Crypto is required only with `STRATA_ENABLE_CONVERSATION_DISK=ON`. This remains single-GPU, matching the released RAM cache.
Validation at `1b26ecf` (the final commit only adds Windows setup documentation and the v2 ignore rule):
| Check | Result |
|---|---|
| Linux CUDA builds, disk support ON and OFF | PASS; OFF has no libcrypto dependency and rejects activation before cache-directory I/O |
| Cache-disabled comparison with untouched v0.1.30 | PASS, 72 requests: release / disk-enabled build / disk-disabled build × INT8 batched / K8V4 batched / INT8 ring/spec4; identical tokens, all nine main-state fingerprints, reuse and known answers |
| Six-phase disk lifecycle, INT8 batched / K8V4 batched / wrapped INT8 ring | PASS ×3: eviction-only writes, restart reuse, draft read-back, quota bounds, admission denial, changed-tokenizer and corruption fallback |
| Long disk restart | 68,571 prompt tokens restored; draft read-back covers 68,572 cells with only 32,840 resident |
| GPU snapshot fixture | PASS, 1,781 checks |
| Long RAM growth/rewind with spec4; RAM off/on parity; HTTP reuse/eviction/cancellation/recovery | PASS ×3 |
| Linux C++ host tests | PASS: codec, store (including fault injection), cache, memory, checkpoint policy, validation and transfer; the five CPU-only suites also pass ASan/UBSan |
| Python | 41 cache-harness tests pass normally and with `-O`; 53 server tests pass |
| CLI | 26 invalid-argument/configuration cases rejected before cache-directory I/O |
| Windows host build and tests | PASS, MSVC / OpenSSL 3.6.4 on `windows-2022`: all five CPU-only C++ suites; [run and logs](https://github.com/jeremiahritchey/Strata/actions/runs/36763048674) |
Linux model testing used one RTX 4090, IQ3_S, CUDA 13.4, CPUs 0–31 with NUMA interleaving, fixed expert residency, and diagnostic state/draft read-back. All ten GPU/model jobs passed. These are correctness gates, not throughput measurements. The Windows workflow ran on a separate validation branch whose only addition to the tested code is that workflow; it is not part of this PR.
**Remaining platform gate:** Windows CUDA engine/model lifecycle execution is still needed. The hosted Windows run verifies OpenSSL/MSVC and CPU/file-store behavior, not GPU restore. `docs/DETAILS.md` now includes Windows OpenSSL discovery, build and runtime-DLL requirements.
Full asset hashing adds startup I/O; path/settings binding is intentionally strict. Disk files contain conversation content. Older draft v1 snapshots are not reused and their directory is left untouched. Windows GPU/HIP execution, power-loss survival and full-model weight-swap/cross-quant cache exchange are not established by the local tests. No multi-GPU disk support, page deduplication, planning files or general Pi benchmark tooling is included.
The streaming-envelope approach follows Marmaduke Woodman's (@maedoc) NVMe work in #52 (`6648be7`), adapted to the shared snapshot core. The RAM admission work from @midhatn is in the released base; the identity diagnostics address the owner/@QilinWan review in #57.
<details>
<summary>Reproduction (Linux, CUDA 13.4, RTX 4090, IQ3_S)</summary>
Build with the setup-pinned llama.cpp checkout and OpenSSL development headers available:
```sh
cmake -S . -B build-review-cuda -DCMAKE_BUILD_TYPE=Release \
-DSTRATA_ENABLE_CUDA=ON -DSTRATA_ENABLE_HIP=OFF \
-DCMAKE_CUDA_COMPILER=/usr/local/cuda-13.4/bin/nvcc \
-DCMAKE_CUDA_ARCHITECTURES=89 -DSTRATA_PORTABLE=ON \
-DSTRATA_NATIVE_EXPERTS=ON -DSTRATA_PREFILL_MMQ=OFF \
-DSTRATA_BUILD_TESTS=OFF -DSTRATA_BUILD_CONVERSATION_TESTS=ON \
-DSTRATA_ENABLE_CONVERSATION_DISK=ON \
-DSTRATA_GGML_DIR=/absolute/path/to/llama.cpp
cmake --build build-review-cuda -j 2 --target strata conversation_snapshot_test \
conversation_file_test conversation_store_test conversation_cache_test \
conversation_memory_test conversation_validation_test conversation_transfer_test conv_cache_test
ctest --test-dir build-review-cuda --output-on-failure \
-R '^(conversation_(file|store|cache|memory|snapshot|validation|transfer)_test|conv_cache_test)$'
python -m unittest discover -s tools -p 'test_conversation_cache*.py'
python -O -m unittest discover -s tools -p 'test_conversation_cache*.py'
python -m unittest serve.test_server
```
GPU/model tests require an exclusive GPU window. Use absolute asset paths in the engine configs, a fixed `--expert-cache 6000` and `--pcie-frac 0.55`. The tested modes are:
| Config | Context | KV | Resident cells |
|---|---:|---|---:|
| `int8-batched.json` | 16384 | int8 | 0 |
| `k8v4-batched.json` | 16384 | k8v4 | 0 |
| `int8-ring.json` | 131072 | int8 | 32768 |
Each output directory must be new. Each lifecycle starts six private engines and changes only copied tokenizer files/generated snapshots beneath that directory:
```sh
mkdir -p logs/pr190-repro
export STRATA_SNAPSHOT_VERIFY=1
export STRATA_MTP_BATCH=1
for mode in int8-batched k8v4-batched; do
numactl --interleave=all --physcpubind=0-31 python tools/conversation_cache_disk.py \
--config "$mode.json" --engine build-review-cuda/strata \
--draft-path batched --output "logs/pr190-repro/$mode" --run
done
numactl --interleave=all --physcpubind=0-31 python tools/conversation_cache_disk.py \
--config int8-ring.json --engine build-review-cuda/strata \
--draft-path ring --paragraphs 4096 --output logs/pr190-repro/int8-ring --run
```
The lifecycle gate requires actual batched/ring evidence, restart draft read-back, exact output/main-state parity, known answers, quota bounds, cold fallback and the named tokenizer-identity rejection. `docs/DETAILS.md` documents the serving configuration and persistence limits.
</details>
No site
Links install, modelos, releases.