Pull requests / #456

#456 Monitor: the conversation cache card (parked conversations from the engine's CACHE lines)

closed · @sebastianrcnt · 0 コメント · GitHub で見る

Server & APIAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

本文

With `--conversation-cache-mib` the engine parks up to N conversations in RAM (#189), but the web app showed nothing about them. Only the per-request "Reused" column hints that a parked conversation was mounted.

**What this adds**

- **Engine:** `CACHE <live tokens> <parked bytes> <evictions> <superseded> [<tokens>:<bytes> ...]` (parked conversations least recently active first), printed after `READY` and after every `DONE`, only when the conversation cache is on. `ConversationCache` gets a read-only `entries()` accessor; nothing else in the engine changes.
- **Server:** `StrataEngine._pump` takes `CACHE` lines off the stdout stream (so a request's reader never sees them) and keeps the last one; `GET /metrics` returns it as `conversations` (budget and slots from the engine arguments, plus the reported state).
- **Monitor:** a "Conversation cache" card - RAM used of the budget, the live conversation, each slot (tokens, RAM, empty, and which one is evicted next), and the evicted / superseded (#342) counts since the engine started. Hidden without `--conversation-cache-mib`.

**Compatibility:** an older server skips the unknown line (the request loop ignores lines it does not know); with an older engine the card says it is waiting for the first report.

**Tested**

- `python -m unittest serve.test_server`: 89 tests OK, including a new `test_conversation_cache_lines` (the pump keeps CACHE lines out of the request queue, the last one wins, a malformed one changes nothing, `/metrics` returns it).
- Linux, RTX 3090, CUDA 13.2 build of this branch, OrcaRouter IQ3_XXS at 262K with `--conversation-cache-mib 8192 --conversation-cache-slots 4`: three alternating conversations plus a follow-up turn; the card showed the live conversation (1,919 tokens) and three parked ones (1,753 / 1,106 / 570 tokens, 0.7 of 8.0 GB), matching the engine log.
- Not tested on Windows or HIP; the engine change is a `printf` in the serve loop, nothing platform-specific.

関連リンク

インストール・モデル・リリースへの站内リンク。