Pull requests / #1730

#1730 HIP: experimental Windows gfx906 compatibility with HIP 5.7

open · @tuandat3019 · 0 コメント · GitHub で見る

Setup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAWindowsLinux

本文

Native Windows HIP 5.7 source builds of the opt-in gfx906 backend hit API/host-header incompatibilities, and the prompt activation flag was not addressable through `hipMemcpyToSymbol` with anonymous device linkage on the tested PAL stack. This adds an experimental manual build recipe and the compatibility fixes needed to build and exercise the engine on a single MI50 16 GiB.

Closes #1728.

Changes:
- Use legacy hipBLAS 5.7 GemmEx datatypes while retaining the existing HIP 6+ path; add missing stream-priority aliases and the Windows host wave64 header fallback.
- Use the existing physical-RAM pinned-memory fallback when legacy HIP properties lack a DXGI LUID. Limit the clang 17 / recent MSVC STL C++17 workaround to the Windows gfx906 helper translation unit.
- Give the host-written activation flag external device linkage on Windows gfx906. Add repeated BF16/FP16 toggle coverage to the existing GEMM parity probe, and register gfx906 GEMM / mapped-memory diagnostic targets.
- Report gfx906 wave64 accurately and reject a wave32 GPU before its kernels launch. Document the exact toolchain, private SDK repairs, runtime DLL/library layout, build commands and limits.

The fixes already merged in #808 and #1396 are not repeated. This does not add a performance kernel, rocBLAS dispatch table, setup detection, ready-made Windows engine, vision or multi-GPU support.

Validation on 2026-10-09:
- Upstream base `fb58e0dbc8399662c0e47c76578c6e878b14f6cf`; exact pinned ggml revision `3cf03257f219afbe7334045ff7c6a06ac68c627d`.
- Windows 11; Instinct MI50 16 GiB, reported as AMD Radeon Pro VII, `gfx906:sramecc-:xnack-`, wave64; PRO 26.Q1 driver; native HIP/hipBLAS/rocBLAS 5.7, HIP clang 17, host clang 23 + MSVC 14.44, code object V5. No ZLUDA. The private SDK packaging/header repairs are explicitly documented; an untouched SDK was not tested.
- Release build passed for `strata`, `strata-device`, `hip_gfx906_mapped_alias`, `hip_gfx906_gemm_f16_io_parity`, `iq_multi_parity`, `mmvq_multi_parity` and `sampler_parity`.
- Newly built executables passed device enumeration, arena selftest, mapped-memory kernel read, BF16/FP16 activation toggle, all six GEMM cases, IQ multi parity, MMVQ multi parity and sampler selftest. Five CTest entries were confirmed registered; the GPU checks themselves were invoked directly.
- GEMM relative L2: 0.000206-0.000217 at beta=0, 3.87e-7 at beta=1. IQ and sampler: zero failures. MMVQ: 148,480 outputs, zero bitwise differences.
- Negative architecture check passed: selecting the RX 6800 returned code 1 with `this engine targets gfx906 wave64` before the arena selftest.
- Static PE inspection found `amdhip64.dll` / `hipblas.dll` imports and embedded gfx906 V5 objects; no CUDA runtime imports. Built engine SHA256: `c793595b19b35f2f22ccae0d629853d6494f897da8f1bd575182708f55634f68`.
- `git diff --check` passed. Comparing the changed core/prefill translation units after preprocessing without includes under CUDA and wave32 HIP defines gave unchanged conditional tokens. This is a limited source check, not a backend build or binary identity claim.

**Draft / not ready for review:** this host lacks CUDA and SYCL toolchains and a Linux gfx906 environment. CUDA, wave32 HIP, Linux gfx906 and SYCL still need builds before review, as required by AGENTS.md. No end-to-end model inference or performance A/B of this new binary is claimed. The checks used an isolated build and the shared GPU lease; the production engine/config were not replaced.

関連リンク

インストール・モデル・リリースへの站内リンク。