Pull requests / #898
#898 perf: Q8 resident adaptation (+24.5–28.2% rotation gain on RTX PRO 6000)
closed · @CC-David-CC · 0 commentaires · Sur GitHub
BenchmarksNVIDIA / CUDASecurityDocumentation
Description
Draft integration: **Q8 on RTX PRO 6000 Blackwell 96 GB**, 32K input + 1K output, MTP T4, FP16 KV. Rotation OFF → ON with async + duplex + per-layer admission fixed: **113.224 → 140.979 tok/s (+24.513%)**, reverse pair **111.568 → 142.975 (+28.151%)**. Both pairs have identical output IDs and recorded work. Full-request improvement is 2.88–3.19%; PCIe bytes unchanged, 18.0 GB host-copy payload removed. Fresh upstream comparison: **main + Q8 reader 106.1–106.8 tok/s → full stack 141.7–141.8 tok/s**. The reader is required to load this Q8 PLE; the full configurations have different output/work trajectories. See the report for both paired results and full request times. Depends on #864, #865 and **Hardin22's #876 (author preserved)** plus duplex/integration work. All flags opt-in; presets unchanged. Seven logical implementation commits. Three CUDA memcheck fixtures: zero errors; default-off parity and Q2 lifecycle checks passed. Broader Q8 state/cancellation/long-input coverage remains, so please keep this as a draft. [Full results, upstream TPS, limitations and history](https://github.com/CC-David-CC/Strata-a5500/blob/draft/q8-resident-adaptation/docs/Q8_RESIDENT_ADAPTATION.md) [Immutable evidence and token/work verifier](https://github.com/CC-David-CC/Strata-a5500/tree/8c473b947c5076edd39539bc5d69c082b237842b/bench/results/2026-10-05-q8-resident-adaptation)
Sur le site
Liens install, modèles, releases.