Pull requests / #1150

#1150 HIP gfx103x: PR #540's attention kernel with 8 cells per step and DPP lane exchanges (bit-exact, prompt +6-12% on RX 6900 XT)

closed · @xjc10 · 0 コメント · GitHub で見る

BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDA

本文

Re-filed from #849, which GitHub closed when `main` was rewritten; the two commits cherry-pick onto 0.1.40.1 without conflicts. The description and measurements are those of #849: the pre75 attention kernel on HIP wave32 with `NC_ = 8` cells per step (CUDA keeps 2) and DPP lane exchanges instead of LDS; bit-exact against the previous kernel (KL 0, argmax and top-10 agreement 100%); prompt +4.2 to +6.1% on one card, +6-12% in the configurations of the report in `bench/results`.

HIP ctest on the RX 6900 XT at 0.1.40.1 (`82f46a8`), with this change and #1148 (the `hip_prefill_wmma_gemm_parity` skip) applied: 19 tests, 13 pass, 6 skip (the WMMA / hipBLASLt / fused-kernel tests that do not apply to gfx1030), 0 fail.

Developed with an AI coding assistant; every number above was measured on 2x RX 6900 XT (gfx1030, PCIe 4.0 x8 each) / Ryzen 5 5600X, ROCm 10.0.

関連リンク

インストール・モデル・リリースへの站内リンク。