Pull requests / #484
#484 Monitor: show the conversation cache at work (reuse, hits on switches, slots, cache RAM)
closed · @xidus90 · 0 コメント · GitHub で見る
Server & APINVIDIA / CUDAModels & quantsSecurityDocumentationWindows
本文
## Summary The Monitor tab shows the conversation cache (`--conversation-cache-mib`) at work, under **Context fill**: - **Prompt reused (last N)**: the share of prompt tokens that were not read again. - **Hits on switches (last N)**: requests that switched to another conversation and came back from a parked snapshot. - **Parked conversations**: slots in use against `--conversation-cache-slots`. - **Cache RAM**: parked snapshots plus the kept K/V of the active conversation, against `--conversation-cache-mib`, with the evictions so far. Without the conversation cache the block stays hidden, so the Monitor looks as before. N is `--monitor-cache-window N` or `"monitor_cache_window": N` in `strata-<model>.json`; the default is 20. <img width="1400" height="1300" alt="real2" src="https://github.com/user-attachments/assets/ac535599-f627-453b-aaf1-4f3c28b436f8" /> ## What a switch and a hit are The engine decides, from what it found before it parks or restores anything: - **switch**: the request restored a parked snapshot, read its prompt from token 0, or reused only the root checkpoint. The root checkpoint is the system prompt, which every chat of a client shares. - **continuation**: the next turn of the same chat. It usually mounts its turn checkpoint rather than the live state, because a client sends the previous reply back without the thinking the engine generated. It does not count. - **hit**: a switch whose restore brought back more than another conversation's root. A new chat of the same client therefore counts as a miss. A changed steering mode or thinking level counts as a switch, and as a miss when no parked snapshot matches. With #342's superseded-copy rule, one chat no longer fills the slots with its own earlier turns. ## Changes - **engine**: the `DONE` line carries five more integers after #471's prompt tokens read (fields 15 to 19): switched, restored, parked, parked bytes and evictions. They are appended, so an older server reads the rest, and a newer server with an older engine shows "–". The root checkpoint is marked where `root_at` saves it. `INFO` prints the budget in effect, so `--prompt-cache 0` shows as off. - **serve**: `GET /metrics` gains `conversation_cache`, summed over the newest N history rows that carry the fields. - A thinking-budget request keeps its first pass's switch, restore and reused prompt. The wrap-up pass's prompt is longer than the request's, so its own values would misreport the request. - The identity token for "this request's DONE" is read once the request holds the engine. A request queued behind another one can no longer take that one's fields. - Parked slots, bytes and evictions come only from the running engine process (an epoch counted at every restart) and are not shown while the model is unloaded. - `monitor_cache_window` accepts 1 to 500. `null` counts as absent and `20.0` as 20, as for the other keys. A bad value stops the start before the model loads. - **tools**: `conversation_cache_parity.py` checks the new fields in both of its scenarios. In `reuse`, A+, B-again and A+-checkpoint must report a restored switch. In `exchange`, B-again must not report one, because B never parks there. - **web, docs**: the four bars, their tooltips, and a paragraph in DETAILS.md. ## Tests On Windows 11 with an RTX 3090 (sm_86), CUDA 13.3 and the IQ3_S model, on this branch at `d9ab843` (0.1.35): - `serve.test_server`: 113 tests. `tools`: 255 tests (with `STRATA_GGUF_PY` set). The new tests failed before their code, and each rule was checked by mutation. - `conversation_cache_test` (4184 checks), `conversation_validation_test` (972) and `conversation_snapshot_test` (1889, GPU). - Parity `reuse`: pass, with byte-exact main-model state. Parity `exchange --cache-mib 450`: pass. - A thinking chat through the server, with a system prompt and without a system block (`reasoning_effort: medium`): turn 2 is a continuation (`switched 0`), and turn 3 after a side request is a hit. A `reasoning_budget_tokens: 20` request reports `reused 0`. An engine with `--prompt-cache 0` gives `enabled: false`. Rebased onto `main` at `99f3dbd` (now `ad96f29`): the cache fields moved behind #471's field 14, and the server parses them from `f[15]`. On the rebased branch, `serve.test_server` (140 tests), `serve.test_monitor`, `serve.test_lifecycle`, `serve.test_security` and `tools.test_conversation_cache_parity` pass. The engine has not been rebuilt or measured again since the rebase: the numbers above are from `d9ab843`. For the `exchange` gate on this engine, A+'s snapshot is now estimated at about 423 MiB, against 384 MB on 0.1.31. The 400 MiB from #189's repro therefore no longer exercises the exchange: use at least 450.
関連リンク
インストール・モデル・リリースへの站内リンク。