Pull requests / #262

#262 HIP: gfx1201 (Radeon AI PRO R9700) support, faster packed byte intrinsics, and a gfx1201 hipBLASLt table

closed · @ttio2tech · 0 comentarios · En GitHub

BenchmarksSetup & installAMD / HIPModels & quants

Descripción

Adds support for gfx1201 (RDNA4, wave32) alongside gfx1100 and speeds up the HIP packed-byte paths.

- The build, setup and device check accept gfx1201 wave32 beside gfx1100.
- `__dp4a` used `v_dot4` only on gfx1100; gfx1201 fell back to a per-byte loop in every decode GEMV.
- HIP's `__byte_perm` indexes a byte array in scratch memory; the IQ4_XS/Q2_0 table lookups now use `v_perm_b32` (IQ4_XS attn_qkv 229 -> 34 us, Q2_0 attn_q 95 -> 23 us on the R9700).
- `__vsub4`, `__vsubss4` and `__vcmpne4` work on all four lanes at once instead of a per-lane loop (Q3_K 14-24%).
- `tests/hip/intrinsics` checks `__byte_perm` and the packed byte ops on 1,024 random operands.
- `tools/hip/gfx1201-hipblaslt-100401.txt`: dense prompt GEMM solutions measured on the R9700.

Q2_0 decode: 58 -> 104 tok/s on the R9700. The output of a greedy CLI run is bit-identical before and after the intrinsic changes.

En el sitio

Enlaces a install, modelos, releases.