Pull requests / #323

#323 hip: add experimental RX 5500 XT 8 GB (gfx1012) support

closed · @CC-David-CC · 0 评论 · 在 GitHub 查看

BenchmarksSetup & installAMD / HIPModels & quantsDocumentationWindowsLinux

描述

This adds a manual Linux HIP source-build path for the consumer AMD Radeon RX 5500 XT 8 GB (RDNA1 / gfx1012).

The changes cover RDNA1 signed-byte dot and wave synchronization compatibility, legacy hipBLAS/rocBLAS compatibility, serving without MTP, and omission of unavailable PCIe expert paths. They include targeted GPU tests, a reproducible benchmark, and the hardware configuration and results. The pinned llama.cpp/ggml dependency is unchanged. Upstream README content is preserved, with a short link to the new hardware documentation.

Tested hardware: RX 5500 XT 8 GB, Ryzen 5 3600, 56 GB installed DDR4 at 1866 MT/s (about 54.8 GiB usable), NVMe SSD, Ubuntu 24.04, HIP 5.7.1 and clang 17. Model: Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S, text only, Q8 KV.

Every request used 8,192 input tokens, 9,216 allocated context tokens, and a maximum of 512 output tokens. Single observations, one request at a time, without prompt KV reuse.

| Task | Output tokens | Generation tok/s, MTP off/on | Total seconds, MTP off/on |
| --- | ---: | ---: | ---: |
| Counting | 512 | 11.35 / 16.81 | 128.96 / 132.55 |
| Coding | 125 | 10.45 / 15.30 | 93.71 / 107.01 |
| Writing | 232 | 10.17 / 10.91 | 104.64 / 120.60 |

MTP accelerated generation, but its additional prefill time increased total completion time for these fresh 8K requests. Total time excludes model startup. Full prefill and effective-throughput results, source/binary hashes, and sanitized raw output are included in docs/AMD_HIP_PERFORMANCE.md and docs/benchmarks/2026-09-30-gfx1012.json.

Validation: 16/16 focused GPU tests and real IQ3_S expert parity checks at layers 0, 8, 24, 40, and 47 passed. All six benchmark requests completed without runtime errors. Both coding outputs passed six functional cases. The writing output exceeded its requested word limit; that is documented.

Validation covers this single-GPU, text-only setup at the stated context. Windows HIP, vision, other architectures, concurrency, and a maximum context limit were not established by these runs. MTP, expert caching, and the inference engine originate in upstream Strata.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。