Pull requests / #39

#39 telemetry: decode-only cache hit rate, prompt checkpoint retention, and dashboard

closed · @code-martin · 0 コメント · GitHub で見る

Server & API

本文

### Summary
Refines MoE cache telemetry to accurately report token generation hit rate (excluding prompt prefill skew), adds multi-turn prefix/system prompt KV cache retention, and provides real-time cache metrics in the web UI.

### Key Changes
- **Decode-Only Hit Rate Baseline**:
  - Captures baseline cache lookups and hits at the exact boundary where prompt processing ends and autoregressive decode begins.
  - Excludes static prompt evaluation from token generation hit rate statistics, reflecting true decode performance.
- **Prefix & System Prompt Checkpoint Retention**:
  - Creates prompt checkpoints at every turn boundary (<|im_start|>) rather than solely at the very last one.
  - Preserves the root checkpoint (checks[0], holding system prompt and initial instructions) during FIFO cache eviction, allowing branching conversations and parallel subagents to immediately reuse the shared prefix.
- **Engine Protocol Enhancements**:
  - Extends the DONE line to output 
eq_hits, 
eq_look, 	otal_swaps, and drift.
  - Adds the STATS command for queryable cumulative telemetry.
- **Web UI & REST API**:
  - Exposes 
ecent_hit_pct, 
ecent_hits, 
ecent_lookups, and cumulative counters in /metrics and /api/stats.
  - Adds an **Expert Cache Hit Rate** progress bar and swap counter to the Monitor tab.
  - Adds a dedicated **Hit rate** column to the Recent requests table in the web dashboard.

### Testing
- Verified live on web dashboard: displays accurate per-request decode hit rates (e.g. 51.2%, 45.6%) and aggregate moving window percentages.
- Verified prompt cache hit behavior with multi-turn chat: turn prefixes and system prompt are retained and reused across turns.

関連リンク

インストール・モデル・リリースへの站内リンク。