Pull requests / #969
#969 serve: Q4 at 297.3 tok/s non-MTP (N=8); MTP +30.4% decode (N=2)
closed · @CC-David-CC · 0 commentaires · Sur GitHub
BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentation
Description
## Draft integration **Q4: 276.2 aggregate decode tok/s at N=8**, reproducing the research result of 275.8. RTX PRO 6000 Blackwell 96 GB, 400 W, FP16 KV; each request reads 32K input and produces 512 tokens. Fresh focused build; cells are decode/effective tok/s: | Requests | MTP | Non-MTP | |---:|---:|---:| | 2 | 232.4 / 34.9 | 178.2 / 33.3 | | 4 | 278.0 / 35.8 | 244.4 / 35.3 | | 8 | 276.2 / 35.7 | 297.3 / 36.1 | N=2: +30.4% decode, +4.7% effective. Single screening runs; non-MTP wins at N=8. Integrates rkcth's #846 with authorship preserved and #947; adds opt-in bounded slot waves and counters. Four focused commits from main; defaults preserved. Passed: CUDA build, 11 frontend tests, 64 requests, 14 exact Q4 MTP/control token pairs, IQ3_S 8/12/16-slot checks. Pending: repeats, GPU sanitizer and cancellation/rollback sweeps. Historical partial-residency token differences remain unclassified. Research continues separately. [Full matrix, graph, proof and reproduction](https://github.com/CC-David-CC/Strata-a5500/blob/contrib/concurrent-mtp-waves/docs/CONCURRENT_SERVING.md)
Sur le site
Liens install, modèles, releases.