Pull requests / #311
#311 HIP: support RDNA2 gfx1030 working and gfx103X is untested.
closed · @xendak · 0 comentários · No GitHub
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
Descrição
## What The experimental HIP backend takes gfx1100 (RDNA3) and gfx1201 (RDNA4). This PR adds RDNA2 gfx1030 (RX 6800 / 6800 XT / 6900 XT / 6950 XT) as an unvalidated target: - `include/strata/hip_compat/intrinsics.hpp`: `dp4a` on gfx1030 / gfx1031 / gfx1032 goes through `__builtin_amdgcn_sdot4`, the plain signed `v_dot4_i32_i8`. RDNA2 has no sudot4 (that is gfx11+); the products are the same signed x signed bytes accumulated modulo 2^32, no clamp. The portable fallback stays for other targets - `cmake/hip_backend.cmake`: gfx1030 joins the unvalidated list (the build warns), with a comment on the dot instruction difference - `setup.py`: `AMD_ARCHS` / `AMD_NAMES` / `AMD_CARDS` list the RX 6800 / 6900 series, and `ROCM_INDEXES` gets the `gfx103X-all` TheRock index (not checked to carry the pinned version: a system ROCm 7 is the tested path) - `docs/AMD_HIP.md`: a RDNA2 (gfx1030) section with the hardware report; `README.md` names the cards - `CMakeLists.txt` and `src/kernels/cuda/verify_kernels.cu`: wording only (gfx10.3's wall_clock64 counter is also a constant 100 MHz, in ns) ## Why gfx1030 is wave32 with the same 64 KiB LDS per workgroup that the kernels' tile sizes assume. The one hardware difference is the dot instruction: no `v_dot4_i32_iu8` / sudot4 before gfx11, but the plain signed `v_dot4_i32_i8` gives the same signed x signed byte products modulo 2^32. There is no WMMA, so the QSA scorer takes the same ordered FP32 fallback as gfx1100. ROCm's hipBLASLt ships no gfx1030 kernels, so there is no tuning table and the plain hipBLAS path runs. CMake keeps the arch unvalidated (the build warns) until a maintainer has run it, the same treatment as gfx1101 / gfx1102 / gfx1200. ## Testing On a community machine: **RX 6900 XT 16 GB (gfx1030)**, i7-13700KF (AVX2, no AVX-512), 63 GB RAM, NixOS, system ROCm 7.2.3 from nixpkgs (clang 22, hipBLAS 3.2): - A complete HIP build for gfx1030, made by hand with cmake and ROCm's own `clang++` (the nixpkgs ROCm is not an `/opt/rocm` tree, so setup's `build_engine_hip` was not exercised; the binary was placed in `engine/` for setup to use) - All 28 registered ctest tests pass on the card, including `hip_device_selftest` (on the 0.1.26 build; the speeds below are from an engine 0.1.30 build of this branch, run on the same card) - End to end (Swift 1.5 IQ3_XXS, `--context 131072` with `--kv-resident 32768 --adapt-every 1 --vram-reserve-mib 1024`, 200 greedy tokens): **38-42 tok/s decode with the default 15 CPU pool workers** (8 workers: 35-37; 24 = every logical core: 26-28), consistent even with the 131,072-token context full - Prefill (the same configuration): **246 tok/s on a 2,000-token prompt, 330-339 tok/s at the auto 8,192-token chunk** (7,997 and 15,967 tokens; time to first token 8.9 and 48.3 s); decode after a 16K prefill holds at 45.5 tok/s **Not tested:** gfx1031 / gfx1032 (the same `dp4a` path, no hardware report), setup's own build path and the `gfx103X-all` wheels on gfx1030, images, answer-quality benchmarks.
No site
Links install, modelos, releases.