Pull requests / #1547
#1547 serve: per-layer expert routing metrics for /metrics ("routing_counts", opt-in)
open · @blange48 · 0 comentários · No GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quantsDocumentation
Descrição
## Per-layer expert routing metrics (`"routing_counts"`, opt-in) Which experts do my requests actually use? Today the only way to see that is a `--dump-routing` trace (~4 KB per generated token, rewritten at every start) and an offline script. This PR adds a cheap, always-on view of the same data for `/metrics`, so Prometheus and Grafana can show it. ### What it does - **Engine** (`--routing-counts PATH`, `--routing-counts-every MIN`): one `uint64` counter per (layer, expert) pair, incremented where the routing trace reads the ids (`drive_pool` / `drive_pool_multi`, every layer of every window). Nothing is written per token. The counts are saved as JSON every minute between requests (the same spot as #477's profile save) and on QUIT, through a temporary file renamed over the last one. A 48 × 512 file is ~55-85 KB. - **Server** (`"routing_counts": "<path>"`, optional `"routing_counts_every"`): passes the flags, and `GET /metrics` reads the file (cached on its mtime): - JSON, `experts` section: per layer, the experts seen, the top 32's share, the entropy in bits, and the top 10 experts; - Prometheus text: `strata:expert_routed_total`; per-layer gauges `strata:expert_seen`, `strata:expert_top32_share` and `strata:expert_entropy_bits` (label `layer`); and `strata:expert_routed_by_expert_total` for each layer's 8 most routed experts (labels `layer`, `expert`). That is ~530 series in all, not 24,576. With `topk(10, rate(...))` this gives one curve per active expert in Grafana. - **Without the key** nothing is counted, written or reported. Documented in `docs/DETAILS.md` next to #477. ### Measured 4× RTX 5080 (layer split 12,24,36), Qwen3.8-Flash-Next IQ3_S, `--batch 8 --batch-groups 4`, on a 0.1.40.3-based build. Same config with and without the key, total tok/s at 1 / 4 / 8 clients (T=0.7, 300 tokens, warm second pass): | | 1 | 4 | 8 | |---|---|---|---| | without | 147 | 269 | 394 | | with `routing_counts` | 143 | 268 | 392 | That is within run-to-run noise. Batch slots, tool calls, the conversation cache and the other 124 metric series are unchanged, and it's in our production with a Grafana row on top of it. The counts agree with a `--dump-routing` trace of the same traffic. Layer 0's top experts are 382, 300 and 324 in both, and every layer is counted the same number of times, so no layer is skipped when its experts are all resident. One thing it showed on this model: routing gets more concentrated with depth. The top 32 of 512 experts take ~25-30% of layer 0's routings and ~50-60% in layers 42-47. ### Tests and scope - `serve/test_prometheus.py`: the summary (counts, entropy, an empty layer, a missing file), the gauges and the per-expert counters (8 at most per layer), and the config key → engine flags. `test_prometheus.py` and `test_server.py` pass on this head (270). - The three commits cherry-pick onto 0.1.40.4 without conflicts. The engine part was built and run on 0.1.40.3 (CUDA 13, sm_120); I have not rebuilt the engine on this exact head (0.1.40.4 only touches Pascal kernels). - Counts restart at zero with the engine (a Prometheus counter handles that). The first file appears at the first request a minute after the start. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
No site
Links install, modelos, releases.