Issues / #709
#709 Does the HIP backend work on non-AMD GPUs via chipStar (HIP → SPIR-V)? Spike says yes with a 3-line shim diff
open · @mq-beefcake · 1 commentaires · Sur GitHub
Description
I run Strata on Intel Arc Pro B70 cards (Battlemage, 2×32GB) and found there's currently no Intel path (CUDA needs NVIDIA, HIP needs AMD). I've been experimenting with [chipStar](https://github.com/CHIP-SPV/chipStar) — a HIP→SPIR-V compatibility layer that runs HIP code on Intel GPUs via Level Zero/OpenCL — and it looks like Strata's existing HIP backend may be portable to Intel almost for free. What I did Compiled Strata's kernels with chipStar's `hipcc` (LLVM 22, `--offload=spirv64`, Level Zero backend) on an Arc Pro B70 (BMG-G21, 0xe223), and ran them: - Compiled unmodified: `rope.cu`, `s_gemv.cu`, `bf16_gemv.cu`, and the 2,211-line `native_mmvq.cu` - Only 3 shim changes needed (all in files Strata already treats as backend shims, mirroring the hip_compat pattern): 1. `include/strata/hip_compat/intrinsics.hpp` — `byte_perm()`: guard `__builtin_amdgcn_perm` to gfx targets, add the portable byte-table fallback for everything else 2. same file — `__nanosleep`: `__builtin_amdgcn_s_sleep(1)` only on gfx targets, no-op otherwise 3. `include/strata/kernels/dp4a.hpp` — same `__nanosleep` guard via `__HIP_PLATFORM_SPIRV__` - `__dp4a` goes through `hip_compat`'s existing SWAR fallback (chipStar has no native dp4a) - Correctness: CPU-reference cross-check of `s_gemv_split` output: max relative error 7e-08 (fp32 rounding only) - Benchmarks (B70, warpSize 32 confirmed, 256 units): D2D copy ceiling 1290 GB/s aggregate; `s_gemv_split` tpr=32 on IQ4_NL 4096×2048 = 152 GB/s (untuned first run) What worked out of the box - warpSize 32 (chipStar honors it on Intel via `cl_intel_required_subgroup_size`) - `__shfl_*` (Strata's hip_compat already maps `_sync` variants to non-sync, which chipStar supports) - `__byte_perm`, int `atomicAdd`, dynamic shared memory ≤ 64 KiB (Strata's HIP path already targets that limit, same as Arc) Known gaps - No native dp4a on the chipStar path → SWAR fallback on the MMQ/MMVQ hot path (this is the main open performance question) - `device.cu`'s arch check only accepts gfx architectures — needs an Intel branch if this becomes supported Questions for the maintainers 1. Would you accept a PR that makes the HIP shim SPIR-V-clean (the 3 diffs above are target-guards only; gfx paths unchanged)? That would make the HIP backend compilable for Intel GPUs with no new backend. 2. Is there any interest in "unvalidated third-backend" guidance similar to the gfx1030 note — i.e. community-run builds for Intel Arc via chipStar, with the standard disclaimer? 3. Anything in the runtime device layer (`src/core/device.cu`, memory budgeting) you know would bite on non-AMD HIP runtimes? Happy to share the spike tree and benchmarks. Not requesting official support — mostly asking whether the shim-cleanliness change is upstream-acceptable so the HIP backend stops assuming gfx targets.
Sur le site
Liens install, modèles, releases.