Pull requests / #442
#442 hip: add community gfx1012 support with version-gated legacy compatibility
closed · @CC-David-CC · 0 comments · View on GitHub
BenchmarksMulti-GPUAMD / HIPModels & quantsDocumentationWindows
Description
Updated to upstream 0.1.34 (`1678de3`). Fresh targeted builds and regressions passed; the full fleet/context performance numbers below remain measurements of the prior 0.1.33 source. [Rebase checks](https://github.com/CC-David-CC/Strata-a5500/blob/contrib/gfx1012-community/docs/benchmarks/2026-10-02-upstream-sync-hardware.md) This is the hardware-only split requested in #323, based on upstream 0.1.34 (`1678de3`). The tested card is the **consumer AMD Radeon RX 5500 XT 8GB, gfx1012**. It remains in the community/unvalidated architecture list. ## Changes - Admit gfx1012 for manual HIP builds. - Signed-byte dot via RDNA1 SDWA, adapted from pinned llama.cpp/ggml with MIT attribution preserved. - Version-gate legacy wave synchronization and host-allocation API names; HIP 7 retains its native operations and spellings. - For hipBLAS 0.x, use the real rocBLAS handle/workspace API. Newer hipBLAS retains its existing implementation. - Preserve Q2_0 FP16 signed zero on gfx1012/HIP below 7, with a regression test. Shared serving and performance alternatives are separate contributions. The P4 tests worked using upstream's existing `STRATA_EXPERIMENTAL_SM60`; this branch adds no SM61 flag, Pascal initializer or FP32 GEMM fallback. ## Validation | Check | RX 5500 XT / HIP 5.7.1 | RX 7900 XTX / HIP 7.15 | |---|---|---| | Committed-source build | Pass | Pass | | Intrinsics and wave/shared-memory synchronization | Pass | Pass | | BF16/F16 prefill GEMM against CPU reference | Pass | Pass | | 1024-value signed-zero regression | Pass | Pass | | Real-weight expert parity, layers 0/3/20/47 | Pass | Pass | | Upstream versus hardware-patch token/state identity | Upstream does not admit gfx1012 | Pass, 8 requests per arm | The independent hardware build completed an 8K-input short code request at **15.30 generation tok/s**, with 106.17 s prefill and 114.34 s total for 125 output tokens. Short prose was 11.68 tok/s, 118.14 s total for 218 output tokens. It also completed **8192 input / 2591 output with 131072 total context tokens allocated**, at 15.66 generation tok/s, 139.54 s prefill and 304.98 s total. This uses Q8 KV streaming with 32768 resident cells. It is an allocation test, **not a full 128K-input test**. The combined hardware/serving tree completed full 8K, 32K and 64K input in both MTP modes. These combined measurements are labelled separately. Full 32K/64K input was not repeated on the hardware-only tree, and full 128K input on the 5500 XT remains untested. [Hardware, scope and reproduction](https://github.com/CC-David-CC/Strata-a5500/blob/contrib/gfx1012-community/docs/COMMUNITY_GFX1012_REVIEW.md) | [Measured results](https://github.com/CC-David-CC/Strata-a5500/blob/contrib/gfx1012-community/docs/benchmarks/2026-10-01-community-gfx1012.md) | [Machine-readable evidence](https://github.com/CC-David-CC/Strata-a5500/blob/contrib/gfx1012-community/docs/benchmarks/2026-10-01-community-gfx1012.json) GSQ-RCO IQ3_S, Q8 KV, text only. Other RDNA1 cards, Windows HIP, vision and multi-GPU behavior are untested. The older-host journal contained a workqueue CPU-time warning; no GPU reset or memory-fault entry appeared in that window. Credit: [Niko1221/Strata](https://github.com/Niko1221/Strata) supplies the engine, expert cache and MTP. The RDNA1 dot implementation is adapted from pinned llama.cpp/ggml code.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.