Pull requests / #442

#442 hip: add community gfx1012 support with version-gated legacy compatibility

closed · @CC-David-CC · 0 commentaires · Sur GitHub

BenchmarksMulti-GPUAMD / HIPModels & quantsDocumentationWindows

Description

Updated to upstream 0.1.34 (`1678de3`). Fresh targeted builds and regressions passed; the full fleet/context performance numbers below remain measurements of the prior 0.1.33 source.

[Rebase checks](https://github.com/CC-David-CC/Strata-a5500/blob/contrib/gfx1012-community/docs/benchmarks/2026-10-02-upstream-sync-hardware.md)

This is the hardware-only split requested in #323, based on upstream
0.1.34 (`1678de3`). The tested card is the **consumer AMD Radeon RX 5500 XT
8GB, gfx1012**. It remains in the community/unvalidated architecture list.

## Changes

- Admit gfx1012 for manual HIP builds.
- Signed-byte dot via RDNA1 SDWA, adapted from pinned llama.cpp/ggml with MIT
  attribution preserved.
- Version-gate legacy wave synchronization and host-allocation API names;
  HIP 7 retains its native operations and spellings.
- For hipBLAS 0.x, use the real rocBLAS handle/workspace API. Newer hipBLAS
  retains its existing implementation.
- Preserve Q2_0 FP16 signed zero on gfx1012/HIP below 7, with a regression test.

Shared serving and performance alternatives are separate contributions. The P4
tests worked using upstream's existing `STRATA_EXPERIMENTAL_SM60`; this branch
adds no SM61 flag, Pascal initializer or FP32 GEMM fallback.

## Validation

| Check | RX 5500 XT / HIP 5.7.1 | RX 7900 XTX / HIP 7.15 |
|---|---|---|
| Committed-source build | Pass | Pass |
| Intrinsics and wave/shared-memory synchronization | Pass | Pass |
| BF16/F16 prefill GEMM against CPU reference | Pass | Pass |
| 1024-value signed-zero regression | Pass | Pass |
| Real-weight expert parity, layers 0/3/20/47 | Pass | Pass |
| Upstream versus hardware-patch token/state identity | Upstream does not admit gfx1012 | Pass, 8 requests per arm |

The independent hardware build completed an 8K-input short code request at
**15.30 generation tok/s**, with 106.17 s prefill and 114.34 s total for 125
output tokens. Short prose was 11.68 tok/s, 118.14 s total for 218 output tokens.

It also completed **8192 input / 2591 output with 131072
total context tokens allocated**, at 15.66 generation tok/s,
139.54 s prefill and 304.98 s total. This uses Q8 KV
streaming with 32768 resident cells. It is an allocation test, **not a full
128K-input test**.

The combined hardware/serving tree completed full 8K, 32K and 64K input in both
MTP modes. These combined measurements are labelled separately. Full 32K/64K
input was not repeated on the hardware-only tree, and full 128K input on the
5500 XT remains untested.

[Hardware, scope and reproduction](https://github.com/CC-David-CC/Strata-a5500/blob/contrib/gfx1012-community/docs/COMMUNITY_GFX1012_REVIEW.md)
| [Measured results](https://github.com/CC-David-CC/Strata-a5500/blob/contrib/gfx1012-community/docs/benchmarks/2026-10-01-community-gfx1012.md)
| [Machine-readable evidence](https://github.com/CC-David-CC/Strata-a5500/blob/contrib/gfx1012-community/docs/benchmarks/2026-10-01-community-gfx1012.json)

GSQ-RCO IQ3_S, Q8 KV, text only. Other RDNA1 cards, Windows HIP, vision and
multi-GPU behavior are untested. The older-host journal contained a workqueue
CPU-time warning; no GPU reset or memory-fault entry appeared in that window.

Credit: [Niko1221/Strata](https://github.com/Niko1221/Strata) supplies the engine,
expert cache and MTP. The RDNA1 dot implementation is adapted from pinned
llama.cpp/ggml code.

Sur le site

Liens install, modèles, releases.