Pull requests / #1000
#1000 HIP: gfx1102 (RX 7600) community-validated, with an end-to-end run (follow-up to #192)
closed · @nexus2905 · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installAMD / HIPModels & quantsDocumentation
描述
Follow-up to #192, which was closed pending (1) a startup check against the exact compiled architecture and (2) an end-to-end RX 7600 run.
## (1) Exact-architecture startup check
This is already upstream: `src/core/device.cu` compares the running card's `gcnArchName` (up to the `:` feature suffix) with `STRATA_HIP_ARCHS`, the architectures the binary was compiled for. A gfx1100 build carried to a gfx1102 card stops with the card's name, its architecture and the build's list. This PR doesn't change it.
## Changes
- `cmake/hip_backend.cmake`: move gfx1102 from `_strata_hip_unvalidated` to `_strata_hip_community` (a status message instead of a warning), and update the comment
- `setup.py`: add gfx1102 to `AMD_ARCHS`, `AMD_NAMES` and `AMD_CARDS`, and map it to TheRock's `gfx110X-dgpu` index in `ROCM_INDEXES`. Without that entry, a PC without a system ROCm would hit `KeyError` in `rocm_index()`
- `docs/AMD_HIP.md`: gfx1102 in the title and arch lists, and an RX 7600 entry under *Community-validated cards*
## (2) End-to-end run: RX 7600 8 GB, engine 0.1.39
Ryzen 5 5600 (AVX2, no AVX-512), 64 GiB DDR4, B450 board (PCIe 3.0 x8, 7.0 GB/s probed), Ubuntu 26.04, system ROCm with hipBLASLt 1.4.1 (no gfx1102 tuning table).
- `./setup.sh --backend hip` (with this patch) detects the card, compiles for gfx1102 and prepares the model
- `strata-device --selftest`: OK
- `ctest --test-dir build -E '^(ple_parity|platform_memory_test)$'`: **62/63**. `hip_prompt_attn_wmma` and `hip_prefill_hipblaslt_gemm` were skipped. The one failure is `expert_multi_test`, which requires AVX-512 ("this CPU cannot run the expert kernel"); the engine uses the AVX2 kernels on this CPU
- **IQ2_XS**, 32K context, `--kv int8`, MTP `--spec 4 --spec-min-p 0.5`, `--vram-reserve-mib 2560`: 347 expert slots
- decode: **24.0 tok/s** (three 320-token replies, greedy, thinking off). The same benchmark gave 14.2 tok/s on engine 0.1.27
- prompts (fresh, `cache_n` 0): 4,716 tokens at **59.1 tok/s**, 16,535 tokens at **57.4 tok/s**. The engine chose 256-token chunks ("too few cache slots to borrow")
- answers were correct (arithmetic, code)
### Notes for an 8 GB card that also drives the desktop
- With the auto reserve (700 MiB), only ~0.5 GiB was left once the model loaded. The kernel then rejected GNOME Shell's command submissions (`amdgpu: Not enough memory for command submission`), GNOME Shell aborted and the session ended. That happened three times while working out the reserve. `--vram-reserve-mib 2048`–`2560` is stable here; the desktop itself used ~0.6–1.9 GiB depending on open apps. It might be worth a setup hint for AMD cards of 8 GB or less that also drive a display.
- The **Coder IQ1_M** leaves only 2.34 GiB after its dense weights (IQ2_XS: 3.13 GiB). With a reserve that keeps the desktop alive it gets 0 cache slots, and then `verify: needs the profile-filled VRAM expert tier`, so I couldn't run it safely on this card.
- Not tested: the TheRock wheel path for gfx1102 (I used system ROCm), and a hipBLASLt tuning table.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。