Pull requests / #39
#39 telemetry: decode-only cache hit rate, prompt checkpoint retention, and dashboard
closed · @code-martin · 0 Kommentare · Auf GitHub
Beschreibung
### Summary Refines MoE cache telemetry to accurately report token generation hit rate (excluding prompt prefill skew), adds multi-turn prefix/system prompt KV cache retention, and provides real-time cache metrics in the web UI. ### Key Changes - **Decode-Only Hit Rate Baseline**: - Captures baseline cache lookups and hits at the exact boundary where prompt processing ends and autoregressive decode begins. - Excludes static prompt evaluation from token generation hit rate statistics, reflecting true decode performance. - **Prefix & System Prompt Checkpoint Retention**: - Creates prompt checkpoints at every turn boundary (<|im_start|>) rather than solely at the very last one. - Preserves the root checkpoint (checks[0], holding system prompt and initial instructions) during FIFO cache eviction, allowing branching conversations and parallel subagents to immediately reuse the shared prefix. - **Engine Protocol Enhancements**: - Extends the DONE line to output eq_hits, eq_look, otal_swaps, and drift. - Adds the STATS command for queryable cumulative telemetry. - **Web UI & REST API**: - Exposes ecent_hit_pct, ecent_hits, ecent_lookups, and cumulative counters in /metrics and /api/stats. - Adds an **Expert Cache Hit Rate** progress bar and swap counter to the Monitor tab. - Adds a dedicated **Hit rate** column to the Recent requests table in the web dashboard. ### Testing - Verified live on web dashboard: displays accurate per-request decode hit rates (e.g. 51.2%, 45.6%) and aggregate moving window percentages. - Verified prompt cache hit behavior with multi-turn chat: turn prefixes and system prompt are retained and reused across turns.
Mehr auf der Site
Links zu Install, Modellen, Releases.