Pull requests / #94
#94 Add experimental gfx1100 HIP backend for RX 7900 XTX
closed · @2jztricks · 0 comentarios · En GitHub
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentationLinux
Descripción
Superseded by #121: the complete gfx1100 HIP support package with tuned prefill, bounded expert uploads and fresh benchmark results. Please review the replacement PR. --- Adds an opt-in Linux HIP backend for the RX 7900 XTX (`gfx1100`, wave32), so Strata can execute native quantized experts, prefill and MTP on AMD. CUDA remains the existing separate backend; this does not modify the installer or enable mixed-vendor inference. ## Changes - Select HIP runtime/hipBLAS in CMake and compile the existing CUDA-shaped kernels as HIP. Reject unsupported HIP targets and simultaneous CUDA/HIP builds. - Provide the required runtime/intrinsic compatibility layer, including RDNA3 signed packed-byte dot products. Use a scalar QSA fallback and a GR LDS tile that fits gfx1100. - Preserve separately rounded Q8_K quantization operations on HIP. The existing byte-exact test caught LLVM fusing the multiply with the rounding bias; the fix is a source-local compile option, not a relaxed tolerance. - Make the mmap expert source honor native per-layer blob layouts and exact file sizes. This avoids requiring a huge pinned host allocation for native packs. New packs must be generated with `iq_pack.py --experts-bin`. - Add asynchronous mapped-memory handoff/graph replay, intrinsic, QSA and mmap regression tests, plus eight-token fused GR replay coverage. Build and serving instructions are in `docs/AMD_HIP.md`. ## Validation Tested against main `c1e903310f211e6630780c3bd2038778c071c68d` with the repository-pinned llama.cpp dependency. Host: RX 7900 XTX 24 GiB, Ryzen 9 7950X3D, 64 GiB RAM, Fedora-family Linux, ROCm 7.1.52802 / Clang 20. - Complete HIP Release build; **25/25 selected CTests passed**. The docs give the exact command and exclusions. - Real original GSQ-RCO IQ3_XXS server with mmap, automatic GPU expert cache, shipped expert profile and MTP: arithmetic, generated Python function (three runtime assertions), system-marker recall and 1,170-token batched-prefill recall all passed with normal stop completion in one engine process. - CUDA 13 / GCC 15 / sm_89 executable build passed; GR, sampler, quantization and mmap regression tests passed. CPU-only `strata-plan` build also passed. - Independent static review and `git diff --check` passed. ## Limits Experimental support is restricted to Linux gfx1100/wave32. No HIP ggml MMQ path, automated installer support, vision validation, full-context stress test, general answer-quality equivalence or mixed AMD/NVIDIA execution is claimed. Other AMD architectures need separate validation. `ple_parity` needs an external fixture; `platform_memory_test` requests 256 MiB of locked memory and failed under the test shell's 8 MiB memlock limit. Both are explicitly excluded from the 25-test count. The original model smoke used mechanical RAID0 and incurred cold page faults, so it is compatibility evidence, not a throughput benchmark. Earlier measurements from a larger working branch are intentionally not attributed to this patch. Related to #48.
En el sitio
Enlaces a install, modelos, releases.