Pull requests / #94

#94 Add experimental gfx1100 HIP backend for RX 7900 XTX

closed · @2jztricks · 0 commentaires · Sur GitHub

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentationLinux

Description

Superseded by #121: the complete gfx1100 HIP support package with tuned prefill, bounded expert uploads and fresh benchmark results. Please review the replacement PR.

---

Adds an opt-in Linux HIP backend for the RX 7900 XTX (`gfx1100`, wave32), so Strata can execute native quantized experts, prefill and MTP on AMD. CUDA remains the existing separate backend; this does not modify the installer or enable mixed-vendor inference.

## Changes

- Select HIP runtime/hipBLAS in CMake and compile the existing CUDA-shaped kernels as HIP. Reject unsupported HIP targets and simultaneous CUDA/HIP builds.
- Provide the required runtime/intrinsic compatibility layer, including RDNA3 signed packed-byte dot products. Use a scalar QSA fallback and a GR LDS tile that fits gfx1100.
- Preserve separately rounded Q8_K quantization operations on HIP. The existing byte-exact test caught LLVM fusing the multiply with the rounding bias; the fix is a source-local compile option, not a relaxed tolerance.
- Make the mmap expert source honor native per-layer blob layouts and exact file sizes. This avoids requiring a huge pinned host allocation for native packs. New packs must be generated with `iq_pack.py --experts-bin`.
- Add asynchronous mapped-memory handoff/graph replay, intrinsic, QSA and mmap regression tests, plus eight-token fused GR replay coverage. Build and serving instructions are in `docs/AMD_HIP.md`.

## Validation

Tested against main `c1e903310f211e6630780c3bd2038778c071c68d` with the repository-pinned llama.cpp dependency. Host: RX 7900 XTX 24 GiB, Ryzen 9 7950X3D, 64 GiB RAM, Fedora-family Linux, ROCm 7.1.52802 / Clang 20.

- Complete HIP Release build; **25/25 selected CTests passed**. The docs give the exact command and exclusions.
- Real original GSQ-RCO IQ3_XXS server with mmap, automatic GPU expert cache, shipped expert profile and MTP: arithmetic, generated Python function (three runtime assertions), system-marker recall and 1,170-token batched-prefill recall all passed with normal stop completion in one engine process.
- CUDA 13 / GCC 15 / sm_89 executable build passed; GR, sampler, quantization and mmap regression tests passed. CPU-only `strata-plan` build also passed.
- Independent static review and `git diff --check` passed.

## Limits

Experimental support is restricted to Linux gfx1100/wave32. No HIP ggml MMQ path, automated installer support, vision validation, full-context stress test, general answer-quality equivalence or mixed AMD/NVIDIA execution is claimed. Other AMD architectures need separate validation.

`ple_parity` needs an external fixture; `platform_memory_test` requests 256 MiB of locked memory and failed under the test shell's 8 MiB memlock limit. Both are explicitly excluded from the 25-test count. The original model smoke used mechanical RAID0 and incurred cold page faults, so it is compatibility evidence, not a throughput benchmark. Earlier measurements from a larger working branch are intentionally not attributed to this patch.

Related to #48.

Sur le site

Liens install, modèles, releases.