Pull requests / #262
#262 HIP: gfx1201 (Radeon AI PRO R9700) support, faster packed byte intrinsics, and a gfx1201 hipBLASLt table
closed · @ttio2tech · 0 コメント · GitHub で見る
BenchmarksSetup & installAMD / HIPModels & quants
本文
Adds support for gfx1201 (RDNA4, wave32) alongside gfx1100 and speeds up the HIP packed-byte paths. - The build, setup and device check accept gfx1201 wave32 beside gfx1100. - `__dp4a` used `v_dot4` only on gfx1100; gfx1201 fell back to a per-byte loop in every decode GEMV. - HIP's `__byte_perm` indexes a byte array in scratch memory; the IQ4_XS/Q2_0 table lookups now use `v_perm_b32` (IQ4_XS attn_qkv 229 -> 34 us, Q2_0 attn_q 95 -> 23 us on the R9700). - `__vsub4`, `__vsubss4` and `__vcmpne4` work on all four lanes at once instead of a per-lane loop (Q3_K 14-24%). - `tests/hip/intrinsics` checks `__byte_perm` and the packed byte ops on 1,024 random operands. - `tools/hip/gfx1201-hipblaslt-100401.txt`: dense prompt GEMM solutions measured on the R9700. Q2_0 decode: 58 -> 104 tok/s on the R9700. The output of a greedy CLI run is bit-identical before and after the intrinsic changes.
関連リンク
インストール・モデル・リリースへの站内リンク。