Pull requests / #996

#996 hip: accept gfx1010 and gfx1011 (RDNA1) in the arch gate

closed · @AcckiyGerman · 0 commentaires · Sur GitHub

BenchmarksAMD / HIPModels & quants

Description

Adds `gfx1010` and `gfx1011` to `_strata_hip_unvalidated`. No source changes are needed.

**gfx1010 (Navi 10) — tested.** RX 5700 XT 8 GB, Ryzen 5 7600X, 30 GB RAM, Ubuntu 26.04, ROCm 7.14.0a20260612 (`gfx101X-dgpu` wheels), Coder IQ1_M, 32,768-token context, `--mmap-experts --vram-reserve-mib 1024 --spec 4`. Official community benchmark script, greedy, 256 output tokens, 3 runs, medians:

| prompt tokens | prompt tok/s | decode tok/s | TTFT |
|---:|---:|---:|---:|
| 4,096 | 115.4 | 18.0 | 35.5 s |
| 16,384 | 144.3 | 20.3 | 113.6 s |
| 30,000 | 149.0 | 23.3 | 201.4 s |

61 of 61 ctest pass on the card, 6 of 6 `tools/needle_bench.py` checks at 8k and 30k. The card holds 607 of the model's 12,288 experts (profile-prefilled), decode cache hit 31%.

**gfx1011 (Navi 12, Radeon PRO V520 / Pro 5600M) — untested on hardware.** Same RDNA1 ISA, builds with 0 errors on the same wheels.

Two facts the arch list can't carry: hipBLASLt 1.4.0 refuses all 26 prefill GEMM shapes on gfx1010 (`heuristic_count = 0`), so there is no tuning table; and with an AMD iGPU in the box `hsa_init` segfaults enumerating the pair - below HIP, so `HIP_VISIBLE_DEVICES` cannot reach it - and the engine needs `ROCR_VISIBLE_DEVICES` naming one KFD node. The #442 SDWA `dp4a` turned out not to be needed: portable 162.7/24.3, SDWA 158.3/24.0, 16/16 HIP tests on the portable build.

The full report directory is ready if you want it under `bench/results/`.

Sur le site

Liens install, modèles, releases.