Pull requests / #104

#104 The Monitor's live tok/s is a rate over the last seconds, not the mean since the first token

closed · merged 2026-09-29 · @xyzzing · 0 commentaires · Sur GitHub

BenchmarksServer & APIAMD / HIPLinux

Description

## What

`GET /metrics → live.tok_s` was `generated / (now − first_token)` — the **mean since the first token**, not a rate. Its first sample is `1/elapsed`, so the Speed card, the header pill and the `tok_s` sparkline showed impossible numbers for the first instant of every answer and then undershot for the first second (one token counted over up to a second of elapsed time).

Against a mock engine paced at a known **25.0 tok/s**, sampling `GET /metrics` the way the web app does:

| | first sample | max | settled |
| --- | --- | --- | --- |
| before | **27,235.7 tok/s** | 27,235.7 | 25.2 |
| after | 4.0 | **28.0** | 24.9 |

that is **1092× the engine's own rate** before, 1.12× after.

## What changed

- `live.tok_s` is the rate over the last `RATE_WINDOW_S` (2 s) of `(time, generated)` samples taken per token. While a request is younger than `RATE_MIN_SPAN_S` (0.25 s) there is no rate yet: the mean so far is reported with its span floored there, so no sample can diverge.
- The old formula is kept as `_tok_s_mean()` and is still reported — `live.tok_s_mean`, and `status.tokens_per_s_mean` on `/status`. Nothing is lost.
- A request that ended **without** a `DONE` line (engine death, error, client disconnect) used to record the **previous** request's `generated`/`decode_ms` as its own `decode_tok_s`. The engine's counters are now trusted only when its `DONE` replaced the object captured at request start, and the row also carries the engine's own token count as `engine_generated`.
- Every added field is additive; `serve/web` needed no change.

The engine's own per-request counters — what the Speed card shows for "last request", and the only number to trust for a finished request — are untouched. The server's progress and `done:` log lines keep their whole-request mean on purpose: they describe a finished request, not a live rate.

## Tests

`serve/test_server.py::LiveRate`, no GPU and no pack:

- `test_the_live_number_is_a_rate` — a `MockEngine(delay_s=0.02)` (50 tok/s) with `/metrics` polled at 10 ms: the live feed never exceeds 4× the engine's own rate, and an idle server reports `tok_s: None`.
- `test_a_request_without_a_done_keeps_no_engine_counters` — an engine that dies mid-answer records `decode_tok_s: None` instead of the stale `990.0 tok/s` it used to inherit from the mock's previous request.

Both fail on the unpatched server (the second with exactly `990.0`, the number above) and pass with the change. The rest of the suite is unaffected:

```
python -m unittest serve.test_server -v     # 28 tests, OK
```

## Notes

Found while porting the engine to HIP/ROCm on a Radeon RX 7900 XTX (gfx1100, Linux); this readout was the first thing there that looked wrong, and the fix is platform-independent. The measurements above were taken by replaying a paced `MockEngine` through the unmodified `serve/server.py` and through the patched one.

Sur le site

Liens install, modèles, releases.