Pull requests / #69
#69 telemetry: per-request decode expert cache hit rate in DONE and web monitor
closed · merged 2026-09-28 · @code-martin · 0 评论 · 在 GitHub 查看
描述
### Summary
This PR adds per-request decode expert cache hit rate metrics to the engine DONE line and displays them in the web monitor:
1. **Decode Baseline Isolation**:
- Records cache_hits and lookups baseline at the start of autoregressive decode, excluding prompt prefill lookups from token generation hit rate statistics.
- Computes
eq_hits and
eq_look for the decode phase alone and appends them to the engine's DONE line:
DONE <generated> <prompt> <prompt ms> <decode ms> <finish> <drafts accepted> <drafts offered> <reused> [hits] [lookups]
- Emits a clean log line to stderr:
strata serve: decode expert cache hit rate: 36.8% (5185 / 14107 hits)
2. **Web Monitor Integration**:
- Parses the two optional fields in StrataEngine._parse_done() (backward compatible if omitted).
- Computes hit_rate in Server.run() and records it in request history.
- Adds a **Hit rate** column to the *Recent requests* table in index.html and formats it in pp.js.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。