Pull requests / #898

#898 perf: Q8 resident adaptation (+24.5–28.2% rotation gain on RTX PRO 6000)

closed · @CC-David-CC · 0 コメント · GitHub で見る

BenchmarksNVIDIA / CUDASecurityDocumentation

本文

Draft integration: **Q8 on RTX PRO 6000 Blackwell 96 GB**, 32K input + 1K output, MTP T4, FP16 KV.

Rotation OFF → ON with async + duplex + per-layer admission fixed: **113.224 → 140.979 tok/s (+24.513%)**, reverse pair **111.568 → 142.975 (+28.151%)**. Both pairs have identical output IDs and recorded work. Full-request improvement is 2.88–3.19%; PCIe bytes unchanged, 18.0 GB host-copy payload removed.

Fresh upstream comparison: **main + Q8 reader 106.1–106.8 tok/s → full stack 141.7–141.8 tok/s**. The reader is required to load this Q8 PLE; the full configurations have different output/work trajectories. See the report for both paired results and full request times.

Depends on #864, #865 and **Hardin22's #876 (author preserved)** plus duplex/integration work. All flags opt-in; presets unchanged. Seven logical implementation commits. Three CUDA memcheck fixtures: zero errors; default-off parity and Q2 lifecycle checks passed. Broader Q8 state/cancellation/long-input coverage remains, so please keep this as a draft.

[Full results, upstream TPS, limitations and history](https://github.com/CC-David-CC/Strata-a5500/blob/draft/q8-resident-adaptation/docs/Q8_RESIDENT_ADAPTATION.md)
[Immutable evidence and token/work verifier](https://github.com/CC-David-CC/Strata-a5500/tree/8c473b947c5076edd39539bc5d69c082b237842b/bench/results/2026-10-05-q8-resident-adaptation)

関連リンク

インストール・モデル・リリースへの站内リンク。