Pull requests / #1597
#1597 Conversation cache: --conversation-cache-min-tokens parks only conversations of at least N tokens
open · @willthehuman · 0 comments · View on GitHub
NVIDIA / CUDAModels & quantsDocumentationLinux
Description
## What this adds `--conversation-cache-min-tokens N` parks only conversations of at least N tokens. `0` (the default) keeps today's behaviour, so nothing changes unless you set it. ## Why Parking happens on a **switch**, and every parked conversation costs one of `--conversation-cache-slots` **whatever its length**. A client that interleaves short side requests with one long conversation therefore spends the cache on the side requests, and once it is full, `make_room()` evicts the long one oldest-first — which is exactly the reuse the cache was turned on for. The case that made me write it: an agent (Hermes) issuing side prompts of a few hundred tokens alongside one long conversation. With `--conversation-cache-slots 2` the short requests park themselves, take both slots, and the long conversation is gone by the time its next turn arrives. ## Measured RTX 4070 SUPER 12 GB / 32 GB / Linux / Q2_0, `--conversation-cache-mib 2560 --conversation-cache-slots 2`. One 19,198-token conversation (`L`) interleaved with three ~700-token requests, then `L` sent again with the same prefix: | | what the cache did | `L` re-sent | |---|---|---| | `min 0` | side requests park; `L` evicted (oldest first) | `0 reused + 19198 read in 41,059 ms` (wall 41.4 s) | | `min 12288` | `skip parking (703 tokens, below the 12288 minimum)` ×3 | `19193 reused + 5 read in 135 ms` (wall 0.3 s) | **304x on the prompt read, 41.4 s → 0.3 s wall.** The skip lines are new; the restore line (`restored 19193 tokens (checkpoint) in 68.3 ms`) is the existing one. Three side requests are needed to show this: with two, the warm-up conversation is the one evicted and `L` survives either way. ## What does not change - **The default.** `0` parks every conversation, whatever its size. Any engine started today behaves identically. - **The answers.** The flag only decides which conversations keep a slot. Nothing about a parked conversation's contents changes; a skipped park is a cache miss, not a different computation. - **The other skip paths.** The new check runs before `make_room()`, so a skipped park neither evicts nor stores anything, and it returns the same non-error `true` the byte-budget and RAM-floor paths already take. - **The two enforcement points stay separate.** `wants()` answers for the length alone; the gate is still `enabled() && wants()`, so `--conversation-cache-mib 0` or `--conversation-cache-slots 0` still disables parking outright. ## Plumbed like its neighbours Option struct, usage text, `from_chars` parsing (with `INT32_MAX` as the range, like `--conversation-cache-slots`, and negatives rejected), the construction site, the gate, and the startup `INFO` line — matching what `--conversation-cache-mib` / `-slots` / `-min-free-mib` already do. ## Tests `conversation_cache_test.cpp`, next to the existing `enabled()` block: - the inclusive boundary (12288 parks, 12287 does not) - `0` parks every length, including empty - a negative minimum parks everything, like `0` (the CLI rejects a negative; the predicate should not invert if one arrives another way) - a disabled cache still answers for length, so the two checks stay independent Both faults were then reintroduced by hand to confirm the checks catch them — `>=` → `>` failed on *"the minimum itself and above are parked"*, and unconditional-park failed on *"below the minimum is not parked"*. ## Docs A paragraph in `docs/DETAILS.md` with the log line, the default, and why a skipped park is free. Happy to rename the flag or fold it into an existing one if you would rather it live somewhere else — I picked a separate option because it composes with the slot count and the RAM floor rather than replacing either.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.