Pull requests / #456
#456 Monitor: the conversation cache card (parked conversations from the engine's CACHE lines)
closed · @sebastianrcnt · 0 评论 · 在 GitHub 查看
Server & APIAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux
描述
With `--conversation-cache-mib` the engine parks up to N conversations in RAM (#189), but the web app showed nothing about them. Only the per-request "Reused" column hints that a parked conversation was mounted. **What this adds** - **Engine:** `CACHE <live tokens> <parked bytes> <evictions> <superseded> [<tokens>:<bytes> ...]` (parked conversations least recently active first), printed after `READY` and after every `DONE`, only when the conversation cache is on. `ConversationCache` gets a read-only `entries()` accessor; nothing else in the engine changes. - **Server:** `StrataEngine._pump` takes `CACHE` lines off the stdout stream (so a request's reader never sees them) and keeps the last one; `GET /metrics` returns it as `conversations` (budget and slots from the engine arguments, plus the reported state). - **Monitor:** a "Conversation cache" card - RAM used of the budget, the live conversation, each slot (tokens, RAM, empty, and which one is evicted next), and the evicted / superseded (#342) counts since the engine started. Hidden without `--conversation-cache-mib`. **Compatibility:** an older server skips the unknown line (the request loop ignores lines it does not know); with an older engine the card says it is waiting for the first report. **Tested** - `python -m unittest serve.test_server`: 89 tests OK, including a new `test_conversation_cache_lines` (the pump keeps CACHE lines out of the request queue, the last one wins, a malformed one changes nothing, `/metrics` returns it). - Linux, RTX 3090, CUDA 13.2 build of this branch, OrcaRouter IQ3_XXS at 262K with `--conversation-cache-mib 8192 --conversation-cache-slots 4`: three alternating conversations plus a follow-up turn; the card showed the live conversation (1,919 tokens) and three parked ones (1,753 / 1,106 / 570 tokens, 0.7 of 8.0 GB), matching the engine log. - Not tested on Windows or HIP; the engine change is a `printf` in the serve loop, nothing platform-specific.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。