Pull requests / #969

#969 serve: Q4 at 297.3 tok/s non-MTP (N=8); MTP +30.4% decode (N=2)

closed · @CC-David-CC · 0 评论 · 在 GitHub 查看

BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentation

描述

## Draft integration

**Q4: 276.2 aggregate decode tok/s at N=8**, reproducing the research result of 275.8. RTX PRO 6000 Blackwell 96 GB, 400 W, FP16 KV; each request reads 32K input and produces 512 tokens.

Fresh focused build; cells are decode/effective tok/s:

| Requests | MTP | Non-MTP |
|---:|---:|---:|
| 2 | 232.4 / 34.9 | 178.2 / 33.3 |
| 4 | 278.0 / 35.8 | 244.4 / 35.3 |
| 8 | 276.2 / 35.7 | 297.3 / 36.1 |

N=2: +30.4% decode, +4.7% effective. Single screening runs; non-MTP wins at N=8.

Integrates rkcth's #846 with authorship preserved and #947; adds opt-in bounded slot waves and counters. Four focused commits from main; defaults preserved.

Passed: CUDA build, 11 frontend tests, 64 requests, 14 exact Q4 MTP/control token pairs, IQ3_S 8/12/16-slot checks.

Pending: repeats, GPU sanitizer and cancellation/rollback sweeps. Historical partial-residency token differences remain unclassified. Research continues separately.

[Full matrix, graph, proof and reproduction](https://github.com/CC-David-CC/Strata-a5500/blob/contrib/concurrent-mtp-waves/docs/CONCURRENT_SERVING.md)

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。