Issues / #1347

#1347 serve: a full conversation cache drops the snapshot at the physical-RAM gate instead of evicting to pass it (270k tokens re-read, 96 s, 32 times)

open · @alanthinker · 0 comentários · No GitHub

BenchmarksServer & APINVIDIA / CUDAModels & quantsLinux

Descrição

The physical-RAM admission gate for parking a conversation is final: when it fails, the snapshot is thrown away even though the cache is holding parked conversations whose RAM it could take back. A long prompt then reads its whole context again.

## Environment

- Strata 0.1.40, Linux, single RTX 4080 SUPER (32 GB), 91 GB RAM, i7-13790F
- IQ3_S (GSQ-RCO), `--max-context 524288`, `--kv int8`, `--kv-resident 32768`
- Experts **resident** (the default here): the engine logs `loaded 46.84 GiB` after `cudaHostRegister ... ok`, so 46.84 GiB is pinned and not reclaimable
- `--conversation-cache-mib 25600 --conversation-cache-slots 25 --conversation-cache-min-free-mib 6144` (the defaults in the deploy script)

## What happens

An agent harness keeps several long conversations. Every time one of them parks, the physical-RAM gate refuses, and the next request re-reads the whole prompt.

From one session's log (before the fix), 95 rejections:

```
strata serve: conversation cache: skip parking (physical RAM admission; need 4842 MiB plus 6144 MiB floor, or telemetry unavailable)
strata serve: prompt 276055 tokens = 0 reused + 276055 read in 97676 ms (2826.2 tok/s), ...
strata serve: prompt 276907 tokens = 0 reused + 276907 read in 98931 ms (2799.0 tok/s), ...
strata serve: prompt 273596 tokens = 0 reused + 273596 read in 97121 ms (2817.1 tok/s), ...
```

32 of them were a full re-read of a ~270k-token context at ~2800 tok/s, i.e. **~96 s each**.

## Why it fails

`make_room()` only balances the cache's own budget (`--conversation-cache-mib` and the slots). The second gate then asks the host for real free RAM:

```
additional + --conversation-cache-min-free-mib  <=  /proc/meminfo MemAvailable
```

Measured on this machine:

- a 270k-token snapshot costs **18,410 B/token**, so `additional` is about **4.73 GiB**
- with the default floor that gate wants **10.73 GiB** free
- a busy server had **8.6 GiB** `MemAvailable`

So it fails by ~2.1 GiB — while the cache at that moment held **18.94 GiB** in 12 parked conversations (`parked=12 bytes=20341561640`), i.e. more than four times what the snapshot needed. Evicting one or two of them would have admitted it.

The asymmetry is that `make_room()` frees RAM by budget, and the gate then decides on RAM, but nothing connects the two: the gate's verdict is final and `make_room()` is never asked again.

## Proposed fix

Evict least-recently-active parked conversations until the gate passes, bounded by the slot count. Each `ConversationBuffer` is a list of 16 MiB segments and each segment is its own allocation, far above glibc's mmap threshold, so an evicted entry is back with the kernel by the time the next check reads `/proc/meminfo` — no waiting is needed between the eviction and the re-check.

The floor-only re-check after the capture is left alone: by then the snapshot is already built and evicting cannot make room for it.

Implementation (2 files, ~20 lines), on top of `82f46a8`:

- `include/strata/core/conversation_cache.hpp` — factor `make_room()`'s loop body into `evict_oldest()`, add `slots()` as the loop bound
- `src/program/generate.cpp` — retry the gate after each eviction; the failure line now reports how many were evicted and how many are still parked, and a new line reports the evictions that made the snapshot fit

No data-structure or interface change: the `private:` section is byte-identical, `make_room()` keeps its signature and its behaviour (its three inlined statements became the body of `evict_oldest()`), and both new methods are additions. `conversation_cache_test` passes 4,197 checks including 6 new ones; `conversation_memory_test` (23) and `serve/test_monitor` (10) pass unchanged.

Commit: https://github.com/alanthinker/Strata/commit/c62b540

## Workaround, for anyone hitting this now

Either lower `--conversation-cache-min-free-mib` or lower `--conversation-cache-mib` below what the cache actually holds. Both are a trade: too little floor and the host goes short, too small a budget and the parked conversations evict each other instead. Neither removes the failure, it only moves the threshold — which is what the fix above is for.

## Related

- #658 proposes folding the conversation-cache budget into the resident-experts headroom, which is the planning side of the same RAM question; this one is about what happens at parking time once the plan is already short.

No site

Links install, modelos, releases.