Issues / #641
#641 AMD Instinct MI50 / MI60 / Radeon VII (gfx906, wave64): a working port, numbers, and PRs to upstream it
closed · @JeanP00l · 6 评论 · 在 GitHub 查看
BenchmarksSetup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
描述
I've been running Strata on **2x AMD Instinct MI50 16 GB (gfx906, wave64)** daily since 29.09, as the backend for a coding agent: Coder IQ1_M, 128K context, layer split across the two cards. It runs at 57-59 tok/s decode on a 17K-token agent prompt. For comparison, llama.cpp on the same model and machine does 26.9 tok/s `tg128`. Each upstream release currently has to be merged into my port by hand, so I'd like to upstream it. I split it into small PRs: - **#638** (draft): gfx906 as an opt-in build (`STRATA_HIP_GFX906=ON`). It adds a compat layer, correctness fixes that changed results or hung (FMA contraction of `__fmul_rn`, ggml MMQ getting cc 9.0, `__byte_perm` in scratch memory, the 25 MHz wall clock) and wave64 kernels. Nothing changes with the option off. It builds for CUDA sm_86, HIP gfx1100 and HIP gfx906. - **#639** (draft): with explicit layer-split points, each GPU loads only its own layers' dense weights. +1,800 experts in VRAM on 2x16 GB. Not AMD-specific. - **#640** (draft): `STRATA_ARENA_MMAP=1`, the native-pack expert arena as a read-only mapped file, for machines where the GPUs hold most experts but RAM is small. RAM available on 32 GB: ~1 GB -> 25 GB. It also covers a ROCclr pin-in-place behaviour worth knowing about. - **#637**: server fix. A dead engine was restarted next to its still-alive predecessor, ran out of VRAM, and left the server answering `400 ... context (0)` forever. Questions for the maintainers: 1. **Build option.** Would you keep gfx906 as its own build option, or should it be folded into `STRATA_ENABLE_HIP` / `cmake/hip_backend.cmake` (wave64 next to wave32)? I can do the second if you prefer. The current shape came from getting the CUDA path running unchanged first, with a CUDA warp as half a wavefront. 2. **Setup.** setup.py detection and an automatic build for gfx906 are not included. AMD dropped gfx906 from current ROCm, so the build uses a community ROCm 7.14 image. Is a "build from source" section in `docs/AMD_HIP.md` enough for now? **Also available, not submitted:** a tensor split (each layer's halves on two cards, an own P2P all-reduce at 7.4-7.9 µs per exchange on gfx906). It works, but is not faster than the layer split on these cards for long prompts. **Next:** the drafts become ready once the cleaned branches have re-run on the cards (ctest, a greedy-text comparison with the port), and I'll add a community benchmark in the `docs/COMMUNITY_BENCHMARKS.md` format. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。