Pull requests / #1654
#1654 [WIP] cuda: optional SM75 main-model QSA padded MMA
open · draft · @Unmaple · 0 コメント · GitHub で見る
BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
本文
WIP / 施工中. Not ready for merge or review. Add an optional SM75 QSA decode path using padded FP16 QK/PV MMA with FP32 accumulation. Only explicit main-model verifier calls at T=2-4 can select it with `STRATA_QSA_SM75_MMA=1`; MTP, prefill and the default path retain native attention. Historical actual-T4 QSA prefix latency reductions on RTX 2080 Ti were 8.034 +/- 0.311% at a 78-token prompt and 11.824 +/- 0.497% at 3038 tokens (95% paired Student-t CI, three independent process rounds). These are not production tok/s claims or 16K/80K validation. **Numerical behavior changes.** The candidate is not bit-exact to native; MTP is excluded because its observed errors were larger. Quality validation is pending. See [measurements and checklist](docs/SM75_QSA_MMA_WIP.md). CUDA SM75 translation-unit compilation passed. Kernel cleanup, full engine build/capture, exact-head numerical fixtures, HIP fallback build, longer contexts and end-to-end ablations remain pending.
関連リンク
インストール・モデル・リリースへの站内リンク。