Pull requests / #1339
#1339 cuda: add an opt-in SM120 Q8 T8 QKV MMQ path
open · @CC-David-CC · 0 comments · View on GitHub
BenchmarksNVIDIA / CUDADocumentation
Description
Experimental default-off SM120 Q8_0 T8 QKV path. Warm-cache synthetic stage: 20.495 to 14.330 us (30.1% less time). Bounds/capture/lifetime fixtures and sanitizer passed. Fresh 8K/512 model runs: 69.41 to 80.91 tok/s, but tokens and work differ: no isolated request-speedup or quality-equivalence claim. Other shapes use native fallback. 
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.