Pull requests / #104
#104 The Monitor's live tok/s is a rate over the last seconds, not the mean since the first token
closed · merged 2026-09-29 · @xyzzing · 0 comments · View on GitHub
BenchmarksServer & APIAMD / HIPLinux
Description
## What `GET /metrics → live.tok_s` was `generated / (now − first_token)` — the **mean since the first token**, not a rate. Its first sample is `1/elapsed`, so the Speed card, the header pill and the `tok_s` sparkline showed impossible numbers for the first instant of every answer and then undershot for the first second (one token counted over up to a second of elapsed time). Against a mock engine paced at a known **25.0 tok/s**, sampling `GET /metrics` the way the web app does: | | first sample | max | settled | | --- | --- | --- | --- | | before | **27,235.7 tok/s** | 27,235.7 | 25.2 | | after | 4.0 | **28.0** | 24.9 | that is **1092× the engine's own rate** before, 1.12× after. ## What changed - `live.tok_s` is the rate over the last `RATE_WINDOW_S` (2 s) of `(time, generated)` samples taken per token. While a request is younger than `RATE_MIN_SPAN_S` (0.25 s) there is no rate yet: the mean so far is reported with its span floored there, so no sample can diverge. - The old formula is kept as `_tok_s_mean()` and is still reported — `live.tok_s_mean`, and `status.tokens_per_s_mean` on `/status`. Nothing is lost. - A request that ended **without** a `DONE` line (engine death, error, client disconnect) used to record the **previous** request's `generated`/`decode_ms` as its own `decode_tok_s`. The engine's counters are now trusted only when its `DONE` replaced the object captured at request start, and the row also carries the engine's own token count as `engine_generated`. - Every added field is additive; `serve/web` needed no change. The engine's own per-request counters — what the Speed card shows for "last request", and the only number to trust for a finished request — are untouched. The server's progress and `done:` log lines keep their whole-request mean on purpose: they describe a finished request, not a live rate. ## Tests `serve/test_server.py::LiveRate`, no GPU and no pack: - `test_the_live_number_is_a_rate` — a `MockEngine(delay_s=0.02)` (50 tok/s) with `/metrics` polled at 10 ms: the live feed never exceeds 4× the engine's own rate, and an idle server reports `tok_s: None`. - `test_a_request_without_a_done_keeps_no_engine_counters` — an engine that dies mid-answer records `decode_tok_s: None` instead of the stale `990.0 tok/s` it used to inherit from the mock's previous request. Both fail on the unpatched server (the second with exactly `990.0`, the number above) and pass with the change. The rest of the suite is unaffected: ``` python -m unittest serve.test_server -v # 28 tests, OK ``` ## Notes Found while porting the engine to HIP/ROCm on a Radeon RX 7900 XTX (gfx1100, Linux); this readout was the first thing there that looked wrong, and the fix is platform-independent. The measurements above were taken by replaying a paced `MockEngine` through the unmodified `serve/server.py` and through the patched one.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.