贡献 / #1597

#1597 Conversation cache: --conversation-cache-min-tokens parks only conversations of at least N tokens

open · @willthehuman · 0 评论 · 去 GitHub 看

NVIDIA / CUDAModels & quantsDocumentationLinux

说明

## What this adds

`--conversation-cache-min-tokens N` parks only conversations of at least N tokens. `0` (the default) keeps today's behaviour, so nothing changes unless you set it.

## Why

Parking happens on a **switch**, and every parked conversation costs one of `--conversation-cache-slots` **whatever its length**. A client that interleaves short side requests with one long conversation therefore spends the cache on the side requests, and once it is full, `make_room()` evicts the long one oldest-first — which is exactly the reuse the cache was turned on for.

The case that made me write it: an agent (Hermes) issuing side prompts of a few hundred tokens alongside one long conversation. With `--conversation-cache-slots 2` the short requests park themselves, take both slots, and the long conversation is gone by the time its next turn arrives.

## Measured

RTX 4070 SUPER 12 GB / 32 GB / Linux / Q2_0, `--conversation-cache-mib 2560 --conversation-cache-slots 2`. One 19,198-token conversation (`L`) interleaved with three ~700-token requests, then `L` sent again with the same prefix:

| | what the cache did | `L` re-sent |
|---|---|---|
| `min 0` | side requests park; `L` evicted (oldest first) | `0 reused + 19198 read in 41,059 ms` (wall 41.4 s) |
| `min 12288` | `skip parking (703 tokens, below the 12288 minimum)` ×3 | `19193 reused + 5 read in 135 ms` (wall 0.3 s) |

**304x on the prompt read, 41.4 s → 0.3 s wall.** The skip lines are new; the restore line (`restored 19193 tokens (checkpoint) in 68.3 ms`) is the existing one.

Three side requests are needed to show this: with two, the warm-up conversation is the one evicted and `L` survives either way.

## What does not change

- **The default.** `0` parks every conversation, whatever its size. Any engine started today behaves identically.
- **The answers.** The flag only decides which conversations keep a slot. Nothing about a parked conversation's contents changes; a skipped park is a cache miss, not a different computation.
- **The other skip paths.** The new check runs before `make_room()`, so a skipped park neither evicts nor stores anything, and it returns the same non-error `true` the byte-budget and RAM-floor paths already take.
- **The two enforcement points stay separate.** `wants()` answers for the length alone; the gate is still `enabled() && wants()`, so `--conversation-cache-mib 0` or `--conversation-cache-slots 0` still disables parking outright.

## Plumbed like its neighbours

Option struct, usage text, `from_chars` parsing (with `INT32_MAX` as the range, like `--conversation-cache-slots`, and negatives rejected), the construction site, the gate, and the startup `INFO` line — matching what `--conversation-cache-mib` / `-slots` / `-min-free-mib` already do.

## Tests

`conversation_cache_test.cpp`, next to the existing `enabled()` block:

- the inclusive boundary (12288 parks, 12287 does not)
- `0` parks every length, including empty
- a negative minimum parks everything, like `0` (the CLI rejects a negative; the predicate should not invert if one arrives another way)
- a disabled cache still answers for length, so the two checks stay independent

Both faults were then reintroduced by hand to confirm the checks catch them — `>=` → `>` failed on *"the minimum itself and above are parked"*, and unconditional-park failed on *"below the minimum is not parked"*.

## Docs

A paragraph in `docs/DETAILS.md` with the log line, the default, and why a skipped park is free.

Happy to rename the flag or fold it into an existing one if you would rather it live somewhere else — I picked a separate option because it composes with the slot count and the RAM floor rather than replacing either.

本站相关内容

相关页面的快捷入口。