Pull requests / #1573

#1573 serve: optionally reclaim host caches before SAVE admission

open · @tuandat3019 · 0 comments · View on GitHub

BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsSecurityWindowsLinux

Description

## Summary

Closes #1572.

SAVE can fail its physical-RAM preflight while auxiliary caches hold memory that the session file will not contain. Add `--session-save-reclaim` to let an opted-in SAVE release those caches under measured pressure and retry the existing admission check. The flag is off by default.

**Draft:** `generate.cpp` is shared by CUDA and HIP. HIP builds and runs here; I cannot build CUDA on this machine. Keeping this draft follows AGENTS.md's backend-build requirement before requesting review. SYCL is not changed or validated.

## What changed

- Add a CPU-only reclamation helper: release retained K/V first, then oldest unpinned parked conversations, then unpinned checkpoints other than the one selected for the file. Re-probe available RAM after each release and stop when the original estimate plus floor fits.
- Preserve all explicit pins, checkpoint order, the first tie selected by `session_deepest_checkpoint`, and that checkpoint's existing storage. Run this before `session_save_checkpoints` takes its pointer. Clear a tail-checkpoint marker only if its checkpoint was removed.
- Unknown telemetry or a saturated/unknown allocation estimate evicts nothing. The normal preflight still gates every copy; freed allocation bytes are never treated as an OS memory reservation.
- Add standalone CPU tests and file-I/O regressions; document the option in help and DETAILS.md.

The default branch through SAVE, session format, RESTORE, memory floor and live tokens/images/device state are unchanged. Evictions trade future prefix-cache hits for headroom and remain in effect if a later admission or I/O check fails. This does not guarantee SAVE under arbitrary memory pressure. No Windows-specific trimming, application integration or hardware-specific setting is included.

## Extra Notes

Validation on Windows, RX 6800 16 GiB, 32 GiB system RAM, HIP gfx1030 / Clang 24.0.0, against main `fb58e0dbc8399662c0e47c76578c6e878b14f6cf` (0.1.41):

- Full HIP `strata` build succeeded. `STRATA_BUILD_CONVERSATION_TESTS=ON`; CTest cache, split-failure, memory, file and new reclaim targets: **5/5 passed**.
- Reclaim test: **119 checks**. Covers denied allocations during eviction, every selected position, deepest/pinned/tied selection, pinned parked entries, retained K/V, byte accounting, early admission after each eviction stage, unchanged no-pressure/unknown-size/unknown-telemetry paths, exhausted admission and empty caches.
- File test: **183 checks**, passed normally and with `STRATA_SESSION_BUFFERED=1`. New cases compare complete file bytes with/without reclamation (pinned and ordinary selection), then inject ENOSPC and verify the old file/live state survive and no temporary remains.
- HIP `conversation_snapshot_test`: **3,901 checks passed** on RX 6800. This is the existing device-state fixture, not a long-context model benchmark.
- `python -m unittest serve.test_slots serve.test_security`: **53 run, 1 skipped, no failures**, using mock engines. The initial sandbox run could not access loopback sockets; the reported run used the normal local environment.
- Real model, fresh isolated processes, Swift-1.5 IQ3_XXS / INT8 KV / MTP, 8K context: default SAVE of **1,096 tokens / 252,872,144 bytes**, restart with the option, RESTORE and SAVE again. SHA-256 matches: `2ceba30fdadb437e98bfb01f3cf05e3100fdc3cfcb3a6d2e9014b99858fab4b4`.
- Real refusal case: set `--conversation-cache-min-free-mib 100000000` rather than consuming system RAM. After three turns, SAVE released **2 checkpoints / 236,134,612 host bytes**, then correctly returned `SERR memory` because the floor remained impossible. The prior file was unchanged; the next request answered `READY`, reusing **1,096** tokens. Successful admission after pressure is covered with injected telemetry in CPU tests, not claimed as a real-memory recovery measurement.

Reproduce the CPU suite:

```sh
cmake -S . -B build -DSTRATA_BUILD_CONVERSATION_TESTS=ON
cmake --build build --target session_save_reclaim_test conversation_file_test conversation_cache_test conversation_memory_test conversation_split_failure_test
ctest --test-dir build -R 'session_save_reclaim_test|conversation_(file|cache|memory|split_failure)_test' --output-on-failure
```

For a live check, add `--session-save-reclaim` to the engine config's `args`, keep `--slot-save-path` enabled, and use the existing `/slots/0?action=save|restore` API. An impossible RAM floor safely exercises refusal without allocating a large pressure buffer. Restore the usual floor for real use.

Not tested: CUDA build/runtime, Linux/macOS, SYCL, or performance at long context on this upstream revision. Earlier local long-session results are deliberately not attributed to this PR. Developed and tested with an AI coding assistant.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.