Pull requests / #1573

#1573 serve: optionally reclaim host caches before SAVE admission

open · @tuandat3019 · 0 comentarios · En GitHub

BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsSecurityWindowsLinux

Descripción

## Summary

Closes #1572.

SAVE can fail its physical-RAM preflight while auxiliary caches hold memory that the session file will not contain. Add `--session-save-reclaim` to let an opted-in SAVE release those caches under measured pressure and retry the existing admission check. The flag is off by default.

**Draft:** `generate.cpp` is shared by CUDA and HIP. HIP builds and runs here; I cannot build CUDA on this machine. Keeping this draft follows AGENTS.md's backend-build requirement before requesting review. SYCL is not changed or validated.

## What changed

- Add a CPU-only reclamation helper: release retained K/V first, then oldest unpinned parked conversations, then unpinned checkpoints other than the one selected for the file. Re-probe available RAM after each release and stop when the original estimate plus floor fits.
- Preserve all explicit pins, checkpoint order, the first tie selected by `session_deepest_checkpoint`, and that checkpoint's existing storage. Run this before `session_save_checkpoints` takes its pointer. Clear a tail-checkpoint marker only if its checkpoint was removed.
- Unknown telemetry or a saturated/unknown allocation estimate evicts nothing. The normal preflight still gates every copy; freed allocation bytes are never treated as an OS memory reservation.
- Add standalone CPU tests and file-I/O regressions; document the option in help and DETAILS.md.

The default branch through SAVE, session format, RESTORE, memory floor and live tokens/images/device state are unchanged. Evictions trade future prefix-cache hits for headroom and remain in effect if a later admission or I/O check fails. This does not guarantee SAVE under arbitrary memory pressure. No Windows-specific trimming, application integration or hardware-specific setting is included.

## Extra Notes

Validation on Windows, RX 6800 16 GiB, 32 GiB system RAM, HIP gfx1030 / Clang 24.0.0, against main `fb58e0dbc8399662c0e47c76578c6e878b14f6cf` (0.1.41):

- Full HIP `strata` build succeeded. `STRATA_BUILD_CONVERSATION_TESTS=ON`; CTest cache, split-failure, memory, file and new reclaim targets: **5/5 passed**.
- Reclaim test: **119 checks**. Covers denied allocations during eviction, every selected position, deepest/pinned/tied selection, pinned parked entries, retained K/V, byte accounting, early admission after each eviction stage, unchanged no-pressure/unknown-size/unknown-telemetry paths, exhausted admission and empty caches.
- File test: **183 checks**, passed normally and with `STRATA_SESSION_BUFFERED=1`. New cases compare complete file bytes with/without reclamation (pinned and ordinary selection), then inject ENOSPC and verify the old file/live state survive and no temporary remains.
- HIP `conversation_snapshot_test`: **3,901 checks passed** on RX 6800. This is the existing device-state fixture, not a long-context model benchmark.
- `python -m unittest serve.test_slots serve.test_security`: **53 run, 1 skipped, no failures**, using mock engines. The initial sandbox run could not access loopback sockets; the reported run used the normal local environment.
- Real model, fresh isolated processes, Swift-1.5 IQ3_XXS / INT8 KV / MTP, 8K context: default SAVE of **1,096 tokens / 252,872,144 bytes**, restart with the option, RESTORE and SAVE again. SHA-256 matches: `2ceba30fdadb437e98bfb01f3cf05e3100fdc3cfcb3a6d2e9014b99858fab4b4`.
- Real refusal case: set `--conversation-cache-min-free-mib 100000000` rather than consuming system RAM. After three turns, SAVE released **2 checkpoints / 236,134,612 host bytes**, then correctly returned `SERR memory` because the floor remained impossible. The prior file was unchanged; the next request answered `READY`, reusing **1,096** tokens. Successful admission after pressure is covered with injected telemetry in CPU tests, not claimed as a real-memory recovery measurement.

Reproduce the CPU suite:

```sh
cmake -S . -B build -DSTRATA_BUILD_CONVERSATION_TESTS=ON
cmake --build build --target session_save_reclaim_test conversation_file_test conversation_cache_test conversation_memory_test conversation_split_failure_test
ctest --test-dir build -R 'session_save_reclaim_test|conversation_(file|cache|memory|split_failure)_test' --output-on-failure
```

For a live check, add `--session-save-reclaim` to the engine config's `args`, keep `--slot-save-path` enabled, and use the existing `/slots/0?action=save|restore` API. An impossible RAM floor safely exercises refusal without allocating a large pressure buffer. Restore the usual floor for real use.

Not tested: CUDA build/runtime, Linux/macOS, SYCL, or performance at long context on this upstream revision. Earlier local long-session results are deliberately not attributed to this PR. Developed and tested with an AI coding assistant.

En el sitio

Enlaces a install, modelos, releases.