Pull requests / #1338

#1338 cuda: reuse secondary Q8 copies in native async refills

open · @CC-David-CC · 0 コメント · GitHub で見る

NVIDIA / CUDADocumentation

本文

Depends on perf/q8-compact-fill-main. Default-off native async D2D reuse. RTX PRO 6000 Blackwell, Q8/FP16 KV, 32K/128K input and 1K output: 51.4-64.4% of primary upload payload replaced; initial decode changes +0.94-3.84%. All four token streams match; 32K MTP work differs. One observation per arm. Lifetime/sanitizer and STOP checks passed.

![Results](https://raw.githubusercontent.com/CC-David-CC/Strata-a5500/befb0321d10c0d42ed52efc61c771539c76afb0a/docs/experiments/secondary-refill-main/results.png)

関連リンク

インストール・モデル・リリースへの站内リンク。