Pull requests / #753
#753 serve: isolate concurrent request metrics (depends on #559)
closed · draft · @zengyishou · 0 评论 · 在 GitHub 查看
Server & APIAMD / HIPModels & quantsSecurity
描述
## Dependency and review scope **Draft: blocked on #559. Do not merge before #559.** This branch is based on #559 head `f65301b3bfc621229c7d5ddb1e268f652be43a59`. Because GitHub can only target an existing branch in this repository, the current comparison against `main` also includes #559's unmerged changes. The contribution from this PR is the final single commit only; after #559 lands this branch must be rebased and the diff rechecked before marking ready. Review the isolated increment here: [incremental diff](https://github.com/zengyishou/Strata/compare/f65301b3bfc621229c7d5ddb1e268f652be43a59...fix/batch-request-metrics). The same increment is also proposed directly to #559's author in [the focused companion PR](https://github.com/blange48/Strata/pull/2). These are alternative review routes for the same change, not patches to apply twice. ## Summary Follow-up to Niko1221/Strata#559, prepared on its `multi-sequence-batching` branch. Concurrent requests currently share the engine's last protocol result and mutable live status. One completion can overwrite another request's counters or clear its activity. Bind content-free protocol collectors to each generation iterator operation, distinguish admission DONE from full BDONE, and aggregate continuation segments without double-counting input. Keep request-owned live status, monotonically bucket output rates, and correlate request IDs across HTTP and metric records. Bind delayed slot drains to their original queue and generation; preserve control ownership through cancellation and rejection. This PR does not introduce the batch engine into main and does not include snapshot, ROCm allocation, or fatal-timeout changes. Existing optional API-monitor behavior is unchanged; the new metric collector retains no prompt or answer content. ## Validation Nineteen deterministic regressions cover seven-way reverse completion, interleaved protocol events, admission errors, cancellation, buffered completion, late drains, iterator context isolation and both API dialects. Tests use synthetic inputs and clocks, not private workloads or GPU-performance estimates. Full server suite on macOS/Python 3.12: 222 tests run, 215 passed and 7 skipped. All 25 request-accounting and Prometheus tests passed; JavaScript syntax and diff whitespace checks passed. The existing Prometheus running gauge and per-request latency histograms are preserved.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。