Pull requests / #422

#422 hip: enable experimental Strix Halo gfx1151 source builds

closed · @Donk119 · 0 comentarios · En GitHub

BenchmarksSetup & installAMD / HIPModels & quantsDocumentation

Descripción

## Scope

Enable an explicit, experimental `-DCMAKE_HIP_ARCHITECTURES=gfx1151` source build for Strix Halo. Keep it in the unvalidated architecture list and include gfx1151 in the builtin-guarded signed dot4 path.

This does **not** enable integrated GPUs in setup.py, change memory sizing, install a ROCm runtime, or claim general APU support. Other architectures and defaults are unchanged. Adds architecture-gate regression tests and a manual build/test report.

## Hardware validation

Tested on a native gfx1151 Strix Halo machine with system ROCm 7.2.3 / Clang 22:

- Full HIP build succeeded; 13 Python architecture/setup tests passed.
- HIP intrinsics parity and the 64 MiB device-arena selftest passed.
- CTest plus IQ-fixture rerun: **51 passed, 3 failed for missing reference artifacts, 2 skipped**. Failures: PLE reference pack/block fixtures, and the full Q2_0 expert pack needed by expert_parity/pool_test. Skips: RDNA4-only attention and a GPU-specific hipBLASLt tuning table. Not claiming an all-green suite.
- A real GSQ-RCO Coder IQ1_M invocation completed with exit 0 and answered **Berlin** to the Germany-capital prompt. Decode reported 16 tokens in 2047.8 ms (7.81 tok/s); this is one short low-cache smoke, not a quality/performance benchmark. The diagnostic invocation continued past EOS because `--stop-eos` was not set; that limitation is recorded.

The artifact has 256 experts per layer, so the test used a dimension-matched uncalibrated profile generated by the existing make_profile.py, not the shipped 512-expert profile. Its PLE table is in shard 2. Exact reproduction details are in docs/AMD_STRIX_HALO.md.

## Safety / remaining limits

Existing model workloads had to be temporarily paused: concurrent-load attempts triggered host reclaim/compaction and were aborted by an unchanged external swap watchdog. The successful isolated test retained the watchdog and host-memory floor; services were restored and health-checked. Coexistence, automatic installer sizing, server/API, vision and long contexts are not validated.

AI assistance: local Qwen generated the two small production-code fragments; Hermes/Astra reviewed them and prepared tests/documentation. Build and hardware results came from actual runs on the target machine.

En el sitio

Enlaces a install, modelos, releases.