Pull requests / #192

#192 HIP: support RDNA3 gfx1101/gfx1102 (RX 7600/7700/7800), not only gfx1100

closed · @nexus2905 · 0 comentarios · En GitHub

BenchmarksSetup & installAMD / HIPModels & quantsDocumentation

Descripción

## What

The experimental HIP backend was locked to gfx1100 (RX 7900 XT/XTX) in four places. This PR widens those checks to also accept gfx1101 and gfx1102, the rest of the RDNA3 wave32 family:

- `cmake/hip_backend.cmake`: accept `CMAKE_HIP_ARCHITECTURES` matching `gfx110[012]`
- `src/core/device.cu`: the runtime device check accepts gfx1100/1101/1102 (still requires wave32)
- `include/strata/hip_compat/intrinsics.hpp`: use `__builtin_amdgcn_sudot4` for `dp4a` on gfx1101/gfx1102 too, not just the scalar fallback
- `setup.py`: `AMD_ARCHS` / `AMD_NAMES` list the RX 7700/7800 (gfx1101) and RX 7600 (gfx1102), and the "not supported" messages are updated to match
- `src/core/device_main.cpp`: `strata-device` prints "RDNA3 wave32" instead of a hard-coded gfx1100

Unrelated fix: `CMakeLists.txt` guards the `native_mmvq_multi` test target with `EXISTS`. `bench/micro/native_mmvq_multi.cpp` is missing from the repository, so configuring with `STRATA_BUILD_TESTS=ON` failed.

## Why

gfx1101 and gfx1102 use the same RDNA3 wave32 ISA as gfx1100. They have the same 64 KiB LDS limit that `fused_gr.cu`'s tile size assumes, and they also have the signed dot4 instruction. Nothing in the kernels depends on gfx1100 specifically. The one hardware difference is that gfx1102 has a smaller VGPR file, which affects occupancy, not correctness. The hipBLASLt tuning table in `tools/hip` is keyed by arch, so on these cards setup doesn't find a table and the engine falls back to hipBLASEx, as it already does when a table doesn't match.

## Testing

Tested only on an **RX 7600 (gfx1102, 8 GB)** + Ryzen 5 5600 (AVX2, no AVX-512) + 64 GB RAM, Ubuntu 26.04, system ROCm in `/opt/rocm`:

- `strata` builds with `-DCMAKE_HIP_ARCHITECTURES=gfx1102`, both by hand and through `./setup.sh --backend hip`, which detects the card and compiles the engine
- `strata-device` reports the card correctly and plans an automatically sized expert cache (about 4,100 slots in 8 GB)
- `ctest` (excluding `ple_parity` and `platform_memory_test`, as in docs/AMD_HIP.md): **39/40 pass**, including every HIP parity test (intrinsics, handoff, native QSA, GR, quantize_act, KV, MMQ prefill, ...). `hip_prefill_hipblaslt_gemm` was skipped
- The one failure is `expert_multi_test`: it exercises the AVX-512 CPU expert kernel and exits on a CPU without AVX-512. That isn't related to this change; the engine chooses the AVX2 kernels on such CPUs

**Not tested:** gfx1101 (RX 7700/7800) on real hardware, and a full end-to-end model run on gfx1102. The model download for that is still in progress. I haven't touched docs/AMD_HIP.md; I can update it if you want these cards listed there.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

En el sitio

Enlaces a install, modelos, releases.