Issues / #709

#709 Does the HIP backend work on non-AMD GPUs via chipStar (HIP → SPIR-V)? Spike says yes with a 3-line shim diff

open · @mq-beefcake · 1 评论 · 在 GitHub 查看

AMD / HIPNVIDIA / CUDA

描述

I run Strata on Intel Arc Pro B70 cards (Battlemage, 2×32GB) and found there's currently no Intel path (CUDA needs NVIDIA, HIP needs AMD). I've been experimenting with [chipStar](https://github.com/CHIP-SPV/chipStar) — a HIP→SPIR-V compatibility layer that runs HIP code on Intel GPUs via Level Zero/OpenCL — and it looks like Strata's existing HIP backend may be portable to Intel almost for free.

What I did

Compiled Strata's kernels with chipStar's `hipcc` (LLVM 22, `--offload=spirv64`, Level Zero backend) on an Arc Pro B70 (BMG-G21, 0xe223), and ran them:

- Compiled unmodified: `rope.cu`, `s_gemv.cu`, `bf16_gemv.cu`, and the 2,211-line `native_mmvq.cu`
- Only 3 shim changes needed (all in files Strata already treats as backend shims, mirroring the hip_compat pattern):
  1. `include/strata/hip_compat/intrinsics.hpp` — `byte_perm()`: guard `__builtin_amdgcn_perm` to gfx targets, add the portable byte-table fallback for everything else
  2. same file — `__nanosleep`: `__builtin_amdgcn_s_sleep(1)` only on gfx targets, no-op otherwise
  3. `include/strata/kernels/dp4a.hpp` — same `__nanosleep` guard via `__HIP_PLATFORM_SPIRV__`
- `__dp4a` goes through `hip_compat`'s existing SWAR fallback (chipStar has no native dp4a)
- Correctness: CPU-reference cross-check of `s_gemv_split` output: max relative error 7e-08 (fp32 rounding only)
- Benchmarks (B70, warpSize 32 confirmed, 256 units): D2D copy ceiling 1290 GB/s aggregate; `s_gemv_split` tpr=32 on IQ4_NL 4096×2048 = 152 GB/s (untuned first run)

What worked out of the box

- warpSize 32 (chipStar honors it on Intel via `cl_intel_required_subgroup_size`)
- `__shfl_*` (Strata's hip_compat already maps `_sync` variants to non-sync, which chipStar supports)
- `__byte_perm`, int `atomicAdd`, dynamic shared memory ≤ 64 KiB (Strata's HIP path already targets that limit, same as Arc)

Known gaps

- No native dp4a on the chipStar path → SWAR fallback on the MMQ/MMVQ hot path (this is the main open performance question)
- `device.cu`'s arch check only accepts gfx architectures — needs an Intel branch if this becomes supported

Questions for the maintainers

1. Would you accept a PR that makes the HIP shim SPIR-V-clean (the 3 diffs above are target-guards only; gfx paths unchanged)? That would make the HIP backend compilable for Intel GPUs with no new backend.
2. Is there any interest in "unvalidated third-backend" guidance similar to the gfx1030 note — i.e. community-run builds for Intel Arc via chipStar, with the standard disclaimer?
3. Anything in the runtime device layer (`src/core/device.cu`, memory budgeting) you know would bite on non-AMD HIP runtimes?

Happy to share the spike tree and benchmarks. Not requesting official support — mostly asking whether the shim-cleanliness change is upstream-acceptable so the HIP backend stops assuming gfx targets.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。