Pull requests / #422
#422 hip: enable experimental Strix Halo gfx1151 source builds
closed · @Donk119 · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installAMD / HIPModels & quantsDocumentation
描述
## Scope Enable an explicit, experimental `-DCMAKE_HIP_ARCHITECTURES=gfx1151` source build for Strix Halo. Keep it in the unvalidated architecture list and include gfx1151 in the builtin-guarded signed dot4 path. This does **not** enable integrated GPUs in setup.py, change memory sizing, install a ROCm runtime, or claim general APU support. Other architectures and defaults are unchanged. Adds architecture-gate regression tests and a manual build/test report. ## Hardware validation Tested on a native gfx1151 Strix Halo machine with system ROCm 7.2.3 / Clang 22: - Full HIP build succeeded; 13 Python architecture/setup tests passed. - HIP intrinsics parity and the 64 MiB device-arena selftest passed. - CTest plus IQ-fixture rerun: **51 passed, 3 failed for missing reference artifacts, 2 skipped**. Failures: PLE reference pack/block fixtures, and the full Q2_0 expert pack needed by expert_parity/pool_test. Skips: RDNA4-only attention and a GPU-specific hipBLASLt tuning table. Not claiming an all-green suite. - A real GSQ-RCO Coder IQ1_M invocation completed with exit 0 and answered **Berlin** to the Germany-capital prompt. Decode reported 16 tokens in 2047.8 ms (7.81 tok/s); this is one short low-cache smoke, not a quality/performance benchmark. The diagnostic invocation continued past EOS because `--stop-eos` was not set; that limitation is recorded. The artifact has 256 experts per layer, so the test used a dimension-matched uncalibrated profile generated by the existing make_profile.py, not the shipped 512-expert profile. Its PLE table is in shard 2. Exact reproduction details are in docs/AMD_STRIX_HALO.md. ## Safety / remaining limits Existing model workloads had to be temporarily paused: concurrent-load attempts triggered host reclaim/compaction and were aborted by an unchanged external swap watchdog. The successful isolated test retained the watchdog and host-memory floor; services were restored and health-checked. Coexistence, automatic installer sizing, server/API, vision and long contexts are not validated. AI assistance: local Qwen generated the two small production-code fragments; Hermes/Astra reviewed them and prepared tests/documentation. Build and hardware results came from actual runs on the target machine.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。