Pull requests / #69

#69 telemetry: per-request decode expert cache hit rate in DONE and web monitor

closed · merged 2026-09-28 · @code-martin · 0 commentaires · Sur GitHub

Server & API

Description

### Summary

This PR adds per-request decode expert cache hit rate metrics to the engine DONE line and displays them in the web monitor:

1. **Decode Baseline Isolation**:
   - Records cache_hits and lookups baseline at the start of autoregressive decode, excluding prompt prefill lookups from token generation hit rate statistics.
   - Computes 
eq_hits and 
eq_look for the decode phase alone and appends them to the engine's DONE line:
     DONE <generated> <prompt> <prompt ms> <decode ms> <finish> <drafts accepted> <drafts offered> <reused> [hits] [lookups]
   - Emits a clean log line to stderr:
     strata serve: decode expert cache hit rate: 36.8% (5185 / 14107 hits)

2. **Web Monitor Integration**:
   - Parses the two optional fields in StrataEngine._parse_done() (backward compatible if omitted).
   - Computes hit_rate in Server.run() and records it in request history.
   - Adds a **Hit rate** column to the *Recent requests* table in index.html and formats it in pp.js.

Sur le site

Liens install, modèles, releases.