Pull requests / #793
#793 Pipelined windows over the active slots (+ slot allocator), /metrics for Prometheus (vLLM names) + monitoring kit
closed · @blange48 · 0 Kommentare · Auf GitHub
Setup & installServer & APINVIDIA / CUDAModels & quantsDocumentationWindows
Beschreibung
## Pipelined windows over the active slots (+ slot allocator), `/metrics` for Prometheus These are the parts of #559 that didn't go into 0.1.39, rebased on main (6f32ec0) and adapted to `"parallel"` and its one status per request. Each commit stands on its own. ### 1. Pipelined group windows over the active slots (`src/program/generate.cpp`, 8 lines) 0.1.39 runs the one-GPU batch windows over the active slots. With `--batch-groups`, a group's window still holds every slot of the group, and the idle ones run pad rows through every layer. The server spreads requests over the groups first, so on a 4-group split two requests sit in two windows that are mostly pad rows. Now a group's window holds its slots up to its last active one. This is crazyaimachine's 8119a7f from #559, without the MTP-draft parts. 4 x RTX 5080 (PCIe Gen3), IQ3_S, 262K context, `"parallel": 8`, `--batch-groups 4 --trim-stage-weights`, through the HTTP server, T=0.7, 300 tokens per answer: | Concurrent | 0.1.39 | with this | | ---: | ---: | ---: | | 1 | 122 | 117-122 (unchanged path) | | 2 | 99 | **149** | | 4 | 199 | **222** | | 8 | 366 | 371 | Exactness (`--pcie-frac 0 --adapt-every 1000000`, `STRATA_IQ_MT_MIN=1`): - `tools/batch_test.py`: identical to solo with 8 slots and with 3 (groups partly filled); - `tools/parking_test.py`: identical, with 141 MB of K/V reused on the second park. ### 2. `GET /metrics` for Prometheus, with vLLM's metric names (`serve/prometheus.py`, three commits) - **Format.** A Prometheus scrape (`Accept: text/plain` or openmetrics, or `?format=prometheus`) gets the text format. Anything else, the Monitor tab included, keeps the JSON. - **vLLM names.** Dashboards and alerts written for a vLLM server read Strata unchanged. These are the names current vLLM exports: - `vllm:num_requests_running` / `_waiting` and `vllm:kv_cache_usage_perc`; - the token and request counters; - `vllm:prefix_cache_queries_total` / `_hits_total` (prompt tokens read / reused); - `vllm:spec_decode_num_draft_tokens_total` / `_accepted_tokens_total` (the MTP drafts); - the histograms `vllm:time_to_first_token_seconds`, `vllm:inter_token_latency_seconds` and `vllm:e2e_request_latency_seconds`. - **Strata's own facts.** What vLLM has no name for comes under `strata:`, named after its JSON key: `strata:live_state`, `strata:live_slots`, `strata:live_tok_s`, `strata:last_hit_rate`, `strata:gpu_util` per card, and so on. The Monitor tab and a Grafana panel read the same value. - **Source.** Everything comes from `Service.metrics()` plus three latency histograms observed when a request finishes. With `"parallel"`, the histograms read each request's own status, and running / waiting come from `live.running` / `live.waiting`. - **Tests.** `serve/test_prometheus.py` (6 tests): the text format (one `# TYPE` per family), the histograms, the names against the JSON, and three requests at once against a batch engine double. Documented in `docs/DETAILS.md`. We run it in production behind a Prometheus that also scrapes vLLM servers. The same Grafana panels show both. ### 3. A monitoring kit (`docs/monitoring/`) - `prometheus.yml` and `servicemonitor.yaml` scrape `/metrics`, with the API key sent as a bearer token. - `grafana-strata.json` is a dashboard to import (it asks for the Prometheus data source): - its first row uses only vLLM's names, so it reads a vLLM server too; - its second row shows the engine's `strata:` series. - Every query was checked against a Prometheus scraping the 4-GPU server above. **Not in this PR any more:** the configuration tools. `tools/autoconfig.py` and `tools/tune.py` duplicated what `setup.py` and `tools/calibrate.py` already do. The useful part, measuring the batch settings with every slot busy, is now a patch to `calibrate.py`, posted on #907, which changes the same function. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Mehr auf der Site
Links zu Install, Modellen, Releases.