Pull requests / #484

#484 Monitor: show the conversation cache at work (reuse, hits on switches, slots, cache RAM)

closed · @xidus90 · 0 comentarios · En GitHub

Server & APINVIDIA / CUDAModels & quantsSecurityDocumentationWindows

Descripción

## Summary

The Monitor tab shows the conversation cache (`--conversation-cache-mib`) at work, under **Context fill**:

- **Prompt reused (last N)**: the share of prompt tokens that were not read again.
- **Hits on switches (last N)**: requests that switched to another conversation and came back from a parked snapshot.
- **Parked conversations**: slots in use against `--conversation-cache-slots`.
- **Cache RAM**: parked snapshots plus the kept K/V of the active conversation, against `--conversation-cache-mib`, with the evictions so far.

Without the conversation cache the block stays hidden, so the Monitor looks as before. N is `--monitor-cache-window N` or `"monitor_cache_window": N` in `strata-<model>.json`; the default is 20.


<img width="1400" height="1300" alt="real2" src="https://github.com/user-attachments/assets/ac535599-f627-453b-aaf1-4f3c28b436f8" />


## What a switch and a hit are

The engine decides, from what it found before it parks or restores anything:

- **switch**: the request restored a parked snapshot, read its prompt from token 0, or reused only the root checkpoint. The root checkpoint is the system prompt, which every chat of a client shares.
- **continuation**: the next turn of the same chat. It usually mounts its turn checkpoint rather than the live state, because a client sends the previous reply back without the thinking the engine generated. It does not count.
- **hit**: a switch whose restore brought back more than another conversation's root. A new chat of the same client therefore counts as a miss.

A changed steering mode or thinking level counts as a switch, and as a miss when no parked snapshot matches. With #342's superseded-copy rule, one chat no longer fills the slots with its own earlier turns.

## Changes

- **engine**: the `DONE` line carries five more integers after #471's prompt tokens read (fields 15 to 19): switched, restored, parked, parked bytes and evictions. They are appended, so an older server reads the rest, and a newer server with an older engine shows "–". The root checkpoint is marked where `root_at` saves it. `INFO` prints the budget in effect, so `--prompt-cache 0` shows as off.
- **serve**: `GET /metrics` gains `conversation_cache`, summed over the newest N history rows that carry the fields.
  - A thinking-budget request keeps its first pass's switch, restore and reused prompt. The wrap-up pass's prompt is longer than the request's, so its own values would misreport the request.
  - The identity token for "this request's DONE" is read once the request holds the engine. A request queued behind another one can no longer take that one's fields.
  - Parked slots, bytes and evictions come only from the running engine process (an epoch counted at every restart) and are not shown while the model is unloaded.
  - `monitor_cache_window` accepts 1 to 500. `null` counts as absent and `20.0` as 20, as for the other keys. A bad value stops the start before the model loads.
- **tools**: `conversation_cache_parity.py` checks the new fields in both of its scenarios. In `reuse`, A+, B-again and A+-checkpoint must report a restored switch. In `exchange`, B-again must not report one, because B never parks there.
- **web, docs**: the four bars, their tooltips, and a paragraph in DETAILS.md.

## Tests

On Windows 11 with an RTX 3090 (sm_86), CUDA 13.3 and the IQ3_S model, on this branch at `d9ab843` (0.1.35):

- `serve.test_server`: 113 tests. `tools`: 255 tests (with `STRATA_GGUF_PY` set). The new tests failed before their code, and each rule was checked by mutation.
- `conversation_cache_test` (4184 checks), `conversation_validation_test` (972) and `conversation_snapshot_test` (1889, GPU).
- Parity `reuse`: pass, with byte-exact main-model state. Parity `exchange --cache-mib 450`: pass.
- A thinking chat through the server, with a system prompt and without a system block (`reasoning_effort: medium`): turn 2 is a continuation (`switched 0`), and turn 3 after a side request is a hit. A `reasoning_budget_tokens: 20` request reports `reused 0`. An engine with `--prompt-cache 0` gives `enabled: false`.

Rebased onto `main` at `99f3dbd` (now `ad96f29`): the cache fields moved behind #471's field 14, and the server parses them from `f[15]`. On the rebased branch, `serve.test_server` (140 tests), `serve.test_monitor`, `serve.test_lifecycle`, `serve.test_security` and `tools.test_conversation_cache_parity` pass. The engine has not been rebuilt or measured again since the rebase: the numbers above are from `d9ab843`.

For the `exchange` gate on this engine, A+'s snapshot is now estimated at about 423 MiB, against 384 MB on 0.1.31. The 400 MiB from #189's repro therefore no longer exercises the exchange: use at least 450.

En el sitio

Enlaces a install, modelos, releases.