Pull requests / #1597
#1597 Conversation cache: --conversation-cache-min-tokens parks only conversations of at least N tokens
open · @willthehuman · 0 comentarios · En GitHub
NVIDIA / CUDAModels & quantsDocumentationLinux
Descripción
## What this adds `--conversation-cache-min-tokens N` parks only conversations of at least N tokens. `0` (the default) keeps today's behaviour, so nothing changes unless you set it. ## Why Parking happens on a **switch**, and every parked conversation costs one of `--conversation-cache-slots` **whatever its length**. A client that interleaves short side requests with one long conversation therefore spends the cache on the side requests, and once it is full, `make_room()` evicts the long one oldest-first — which is exactly the reuse the cache was turned on for. The case that made me write it: an agent (Hermes) issuing side prompts of a few hundred tokens alongside one long conversation. With `--conversation-cache-slots 2` the short requests park themselves, take both slots, and the long conversation is gone by the time its next turn arrives. ## Measured RTX 4070 SUPER 12 GB / 32 GB / Linux / Q2_0, `--conversation-cache-mib 2560 --conversation-cache-slots 2`. One 19,198-token conversation (`L`) interleaved with three ~700-token requests, then `L` sent again with the same prefix: | | what the cache did | `L` re-sent | |---|---|---| | `min 0` | side requests park; `L` evicted (oldest first) | `0 reused + 19198 read in 41,059 ms` (wall 41.4 s) | | `min 12288` | `skip parking (703 tokens, below the 12288 minimum)` ×3 | `19193 reused + 5 read in 135 ms` (wall 0.3 s) | **304x on the prompt read, 41.4 s → 0.3 s wall.** The skip lines are new; the restore line (`restored 19193 tokens (checkpoint) in 68.3 ms`) is the existing one. Three side requests are needed to show this: with two, the warm-up conversation is the one evicted and `L` survives either way. ## What does not change - **The default.** `0` parks every conversation, whatever its size. Any engine started today behaves identically. - **The answers.** The flag only decides which conversations keep a slot. Nothing about a parked conversation's contents changes; a skipped park is a cache miss, not a different computation. - **The other skip paths.** The new check runs before `make_room()`, so a skipped park neither evicts nor stores anything, and it returns the same non-error `true` the byte-budget and RAM-floor paths already take. - **The two enforcement points stay separate.** `wants()` answers for the length alone; the gate is still `enabled() && wants()`, so `--conversation-cache-mib 0` or `--conversation-cache-slots 0` still disables parking outright. ## Plumbed like its neighbours Option struct, usage text, `from_chars` parsing (with `INT32_MAX` as the range, like `--conversation-cache-slots`, and negatives rejected), the construction site, the gate, and the startup `INFO` line — matching what `--conversation-cache-mib` / `-slots` / `-min-free-mib` already do. ## Tests `conversation_cache_test.cpp`, next to the existing `enabled()` block: - the inclusive boundary (12288 parks, 12287 does not) - `0` parks every length, including empty - a negative minimum parks everything, like `0` (the CLI rejects a negative; the predicate should not invert if one arrives another way) - a disabled cache still answers for length, so the two checks stay independent Both faults were then reintroduced by hand to confirm the checks catch them — `>=` → `>` failed on *"the minimum itself and above are parked"*, and unconditional-park failed on *"below the minimum is not parked"*. ## Docs A paragraph in `docs/DETAILS.md` with the log line, the default, and why a skipped park is free. Happy to rename the flag or fold it into an existing one if you would rather it live somewhere else — I picked a separate option because it composes with the slot count and the RAM floor rather than replacing either.
En el sitio
Enlaces a install, modelos, releases.