Pull requests / #190

#190 Spill evicted conversation snapshots to a bounded optional disk cache

closed · draft · @jeremiahritchey · 0 Kommentare · Auf GitHub

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentationWindowsLinux

Beschreibung

RAM-evicted conversations can retain reusable state on disk and return after an engine restart. This adds an opt-in, quota-bounded disk tier that consumes the shared snapshot representation and restore path; it does not add another GPU state-application implementation.

**Depends on #189; merge the shared core/RAM change first.** This draft targets `Niko1221/Strata:main`, so its current Files changed view also contains the unmerged core. Review the [NVMe-only comparison](https://github.com/jeremiahritchey/Strata/compare/feat/conversation-cache-unified...feat/conversation-cache-nvme) for the disk additions. After #189 lands, this branch will be updated so the upstream diff contains only the dependent changes.

Refs #57 and #52. This brings the dependent disk review upstream from [the fork draft](https://github.com/jeremiahritchey/Strata/pull/2). General Pi benchmark tooling is excluded.

- Stream a versioned portable envelope with full asset/runtime identity and SHA-256 integrity checking into one admitted image before GPU application.
- Spill only on RAM eviction, with byte/entry quotas, an exclusive process lock, temporary-file cleanup and atomic publication. Disabled persistence performs no filesystem I/O; OpenSSL Crypto is required only for disk-enabled builds.
- Compare active/RAM/disk prefixes, bound promotion staging within the RAM budget, and fall back on read, integrity or admission failures. Rebase persisted checkpoint ages before promotion.

Final 0.1.27 Linux validation passes disk-on/off CUDA builds, 26 CLI rejection cases, host/fault-injection/sanitizer tests, 1,781 GPU snapshot checks and three six-phase model lifecycles: batched INT8, batched K8V4 and wrapped INT8 draft ring. Each lifecycle covers eviction, restart, admission denial, changed-tokenizer identity and corrupted-file fallback with known answers, token/main-state parity and restored draft byte fingerprints. The ring run restores 68,571 tokens across restart, exceeding 32,840 resident draft cells.

Full asset hashing adds cold-start I/O, paths are bound strictly, and whole-image writes occur only on eviction. The latest active turn is therefore not guaranteed durable. The digest cannot name the first mismatching identity field. Disk page-level deduplication is outside this change. Windows/HIP execution, power-loss survival and full-model weight-swap/cross-quantization cache exchange are not established by these tests.

The streaming-envelope approach is adapted from Marmaduke Woodman's (@maedoc) NVMe work in Niko1221/Strata#52 at `6648be7`; the integration uses the shared core and preserves the contributor reference. @midhatn's authored admission contribution remains in the base; @QilinWan's identity/coordinate review is mapped to the documented fields and limits.

The validation summary above applies to the tested 0.1.27 engine. Detailed design/validation notes and raw model/build logs are retained locally; planning documents are excluded from this PR.

Additional default-off regression (2026-09-30): this disk-enabled build, with both caches left at their disabled defaults, matches untouched upstream `a790805` across eight requests in each of three modes: INT8 resident ring/spec-4 (43,170-token prompt), batched INT8, and batched K8V4. The complete upstream/core/NVMe matrix passes 72 requests with exact output-token, comparable main-state, prefix-reuse and known-answer parity. The combined cache harness suite, including disk, default-off and growth negative controls, passes 41 tests. `docs/DETAILS.md` now includes the OpenSSL/build option, config example, quotas, and eviction-only persistence limits.

The shared-core follow-up retains unchanged K/V pages across repeated parking. Disk serialization streams the logical payload independently of RAM segmentation; unused capacity is excluded and reader staging remains bounded. Disk writes still contain whole snapshots and occur only on eviction. The disk-enabled build passes the paired 12-request full/incremental capture gate with INT8 resident-ring/spec-4 at 61,259 prompt tokens, including growth, rewind, exact output/main-state parity and restored draft read-back. See the [independent validation and suggestions](https://github.com/Niko1221/Strata/pull/189#issuecomment-5903850856).

Follow-up validation at `f1b0c16` (core) / `18c475e` (NVMe): all 23 final GPU/model jobs pass, including growth/rewind, tight-budget fallback, cache-on/off parity, isolation, HTTP recovery, the 30-cycle soak, three disk lifecycles, and the 72-request upstream/default-off matrix. Core and NVMe each pass 1,781 GPU snapshot checks. Host sanitizer checks include 4,149 buffer/cache policy checks, 1,452 transfer checks, and 6,732 disk-codec checks. Windows execution remains untested locally.

Portability follow-up (`772e55c`): cache harness text reads/writes explicitly use UTF-8, and isolation fixtures use absolute paths so an engine configured with a different working directory can locate them. This addresses the Windows tester’s two inline findings. All 41 cache harness tests pass under normal and optimized Python. CPU-only reproductions fail on the prior code and pass after the fix for simulated cp1252 defaults and differing engine working directories (image/add/project). These checks do not constitute a Windows engine rerun.

<details>
<summary>Reproduction commands (Linux, IQ3_S / RTX 4090)</summary>

Run from this PR branch's root, with the model installation's Python environment activated. These commands reproduce the Linux RTX 4090 settings (CUDA 13.4, SM89, IQ3_S); replace the two asset paths. The model config must contain absolute asset paths. Model runs and GPU fixtures require an exclusive GPU window. Each model output directory must be new. This is correctness validation, not a throughput benchmark.

```sh
export STRATA_CONFIG=/absolute/path/to/strata-iq3_s.json
export STRATA_LLAMA=/absolute/path/to/llama.cpp
export STRATA_GGUF_PY="$STRATA_LLAMA/gguf-py"
export CUDACXX=/usr/local/cuda-13.4/bin/nvcc
```

The dependency is the setup-pinned llama.cpp commit `3cf03257f219afbe7334045ff7c6a06ac68c627d`. Preserve your config's model/tokenizer/profile paths while selecting the tested context, residency and expert-cache settings:


```sh
python - <<'PY'
import copy
import json
import os
from pathlib import Path

source = json.loads(Path(os.environ['STRATA_CONFIG']).read_text())
keys = {'--max-context', '--kv', '--kv-resident', '--expert-cache', '--pcie-frac'}
args, i = [], 0
while i < len(source['args']):
    flag = source['args'][i]
    if flag in keys or flag.startswith('--conversation-cache-'):
        i += 2
    else:
        args.append(flag)
        i += 1
out = Path('logs/pr-repro')
out.mkdir(parents=True, exist_ok=True)
for name, context, kv, resident in [('int8-ring', 131072, 'int8', 32768),
                                    ('int8-batched', 16384, 'int8', 0),
                                    ('k8v4-batched', 16384, 'k8v4', 0)]:
    cfg = copy.deepcopy(source)
    cfg['args'] = args + ['--max-context', str(context), '--kv', kv,
                          '--kv-resident', str(resident), '--expert-cache', '6000',
                          '--pcie-frac', '0.55']
    cfg['cwd'] = str(Path.cwd())
    cfg.pop('log', None)
    (out / (name + '.json')).write_text(json.dumps(cfg, indent=2) + '\n')
PY
```

CPU-only codec/store checks (OpenSSL development headers and Crypto library required):

```sh
cmake -S . -B build-conversation-host -DSTRATA_ENABLE_CUDA=OFF -DSTRATA_ENABLE_HIP=OFF \
  -DSTRATA_NATIVE_EXPERTS=OFF -DSTRATA_BUILD_TESTS=OFF \
  -DSTRATA_BUILD_CONVERSATION_TESTS=ON -DSTRATA_ENABLE_CONVERSATION_DISK=ON
cmake --build build-conversation-host -j 2 --target conversation_file_test conversation_store_test \
  conversation_cache_test conversation_memory_test conv_cache_test
ctest --test-dir build-conversation-host \
  -R '^(conversation_file_test|conversation_store_test|conversation_cache_test|conversation_memory_test|conv_cache_test)$' --output-on-failure
python -m unittest tools.test_conversation_cache_disk tools.test_conversation_cache_disabled tools.test_conversation_cache_growth
```

Build the disk-enabled CUDA engine using the same settings as the tested core:

```sh
cmake -S . -B build-review-cuda -DCMAKE_BUILD_TYPE=Release \
  -DSTRATA_ENABLE_CUDA=ON -DSTRATA_ENABLE_HIP=OFF -DCMAKE_CUDA_ARCHITECTURES=89 \
  -DSTRATA_PORTABLE=ON -DSTRATA_NATIVE_EXPERTS=ON -DSTRATA_PREFILL_MMQ=OFF \
  -DSTRATA_BUILD_TESTS=OFF -DSTRATA_BUILD_CONVERSATION_TESTS=ON \
  -DSTRATA_ENABLE_CONVERSATION_DISK=ON -DSTRATA_GGML_DIR="$STRATA_LLAMA"
cmake --build build-review-cuda --target strata conversation_snapshot_test -j 2
```

GPU/model checks, only in the exclusive window. Each disk lifecycle starts six private engines and changes only copied tokenizer files and generated snapshots beneath its new output directory:

```sh
ctest --test-dir build-review-cuda -R '^conversation_snapshot_test$' --output-on-failure
export STRATA_SNAPSHOT_VERIFY=1
export STRATA_MTP_BATCH=1
for mode in int8-batched k8v4-batched; do
  numactl --interleave=all --physcpubind=0-31 python tools/conversation_cache_disk.py \
    --config "logs/pr-repro/$mode.json" --engine build-review-cuda/strata \
    --draft-path batched --output "logs/pr-repro/disk-$mode" --run
done
numactl --interleave=all --physcpubind=0-31 python tools/conversation_cache_disk.py \
  --config logs/pr-repro/int8-ring.json --engine build-review-cuda/strata \
  --draft-path ring --paragraphs 4096 --output logs/pr-repro/disk-int8-ring --run
```

The disk tool requires actual batched/ring evidence, restart draft read-back, exact output/main-state parity, known answers, quota bounds and cold fallback on admission/foreign/corrupt inputs. For the separate upstream-versus-disabled comparison, run the reference build and `conversation_cache_disabled.py` commands in #189 from this branch: its candidate is then the disk-enabled binary with both caches left off. Those comparisons use the common upstream hash fields and are correctness checks, not speed measurements.

The growth/rewind commands in #189 also apply to this branch. Disk blobs serialize the logical payload independently of RAM segmentation; file staging includes segment-directory storage. Retaining pages in RAM does not change eviction-only disk writes or introduce disk page deduplication.

</details>

Mehr auf der Site

Links zu Install, Modellen, Releases.