Pull requests / #793

#793 Pipelined windows over the active slots (+ slot allocator), /metrics for Prometheus (vLLM names) + monitoring kit

closed · @blange48 · 0 comments · View on GitHub

Setup & installServer & APINVIDIA / CUDAModels & quantsDocumentationWindows

Description

## Pipelined windows over the active slots (+ slot allocator), `/metrics` for Prometheus

These are the parts of #559 that didn't go into 0.1.39, rebased on main (6f32ec0) and adapted to `"parallel"` and its one status per request. Each commit stands on its own.

### 1. Pipelined group windows over the active slots (`src/program/generate.cpp`, 8 lines)

0.1.39 runs the one-GPU batch windows over the active slots. With `--batch-groups`, a group's window still holds every slot of the group, and the idle ones run pad rows through every layer. The server spreads requests over the groups first, so on a 4-group split two requests sit in two windows that are mostly pad rows. Now a group's window holds its slots up to its last active one. This is crazyaimachine's 8119a7f from #559, without the MTP-draft parts.

4 x RTX 5080 (PCIe Gen3), IQ3_S, 262K context, `"parallel": 8`, `--batch-groups 4 --trim-stage-weights`, through the HTTP server, T=0.7, 300 tokens per answer:

| Concurrent | 0.1.39 | with this |
| ---: | ---: | ---: |
| 1 | 122 | 117-122 (unchanged path) |
| 2 | 99 | **149** |
| 4 | 199 | **222** |
| 8 | 366 | 371 |

Exactness (`--pcie-frac 0 --adapt-every 1000000`, `STRATA_IQ_MT_MIN=1`):
- `tools/batch_test.py`: identical to solo with 8 slots and with 3 (groups partly filled);
- `tools/parking_test.py`: identical, with 141 MB of K/V reused on the second park.

### 2. `GET /metrics` for Prometheus, with vLLM's metric names (`serve/prometheus.py`, three commits)

- **Format.** A Prometheus scrape (`Accept: text/plain` or openmetrics, or `?format=prometheus`) gets the text format. Anything else, the Monitor tab included, keeps the JSON.
- **vLLM names.** Dashboards and alerts written for a vLLM server read Strata unchanged. These are the names current vLLM exports:
  - `vllm:num_requests_running` / `_waiting` and `vllm:kv_cache_usage_perc`;
  - the token and request counters;
  - `vllm:prefix_cache_queries_total` / `_hits_total` (prompt tokens read / reused);
  - `vllm:spec_decode_num_draft_tokens_total` / `_accepted_tokens_total` (the MTP drafts);
  - the histograms `vllm:time_to_first_token_seconds`, `vllm:inter_token_latency_seconds` and `vllm:e2e_request_latency_seconds`.
- **Strata's own facts.** What vLLM has no name for comes under `strata:`, named after its JSON key: `strata:live_state`, `strata:live_slots`, `strata:live_tok_s`, `strata:last_hit_rate`, `strata:gpu_util` per card, and so on. The Monitor tab and a Grafana panel read the same value.
- **Source.** Everything comes from `Service.metrics()` plus three latency histograms observed when a request finishes. With `"parallel"`, the histograms read each request's own status, and running / waiting come from `live.running` / `live.waiting`.
- **Tests.** `serve/test_prometheus.py` (6 tests): the text format (one `# TYPE` per family), the histograms, the names against the JSON, and three requests at once against a batch engine double. Documented in `docs/DETAILS.md`.

We run it in production behind a Prometheus that also scrapes vLLM servers. The same Grafana panels show both.


### 3. A monitoring kit (`docs/monitoring/`)

- `prometheus.yml` and `servicemonitor.yaml` scrape `/metrics`, with the API key sent as a bearer token.
- `grafana-strata.json` is a dashboard to import (it asks for the Prometheus data source):
  - its first row uses only vLLM's names, so it reads a vLLM server too;
  - its second row shows the engine's `strata:` series.
- Every query was checked against a Prometheus scraping the 4-GPU server above.


**Not in this PR any more:** the configuration tools. `tools/autoconfig.py` and `tools/tune.py` duplicated what `setup.py` and `tools/calibrate.py` already do. The useful part, measuring the batch settings with every slot busy, is now a patch to `calibrate.py`, posted on #907, which changes the same function.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.