Pull requests / #895

#895 hip: support Radeon gfx1151 APUs

closed · @DTU-travelpals · 0 comentarios · En GitHub

BenchmarksSetup & installServer & APIAMD / HIPModels & quantsWindowsLinux

Descripción

## Summary

Add Linux HIP support for AMD Radeon 8060S / 8050S APUs (`gfx1151`, RDNA 3.5).

This includes:

- architecture detection and compilation;
- support for AMD’s multi-architecture TheRock wheels;
- shared-memory-aware model selection and expert-cache sizing;
- gfx1151 kernel compatibility;
- setup tests;
- installation and HIP documentation.

The implementation was validated on a Ryzen AI Max+ 395 with Radeon 8060S and 128 GB of unified memory.

## Why gfx1151 needs special handling

The Radeon 8060S is an integrated GPU whose CPU and GPU share the same physical RAM.

On the test machine:

- installed RAM: 128 GB;
- GPU-addressable GTT pool: 116.0 GiB;
- architecture: `gfx1151`;
- wavefront size: 32;
- compute units: 40.

`hipMemGetInfo()` reports the large GPU-addressable GTT pool without subtracting Strata’s ordinary CPU allocations, including its host expert arena. Treating this capacity as independent VRAM can oversubscribe system memory and invoke the OOM killer.

This change:

- recognizes gfx1151 as a supported wave32 HIP architecture;
- labels its GTT capacity as shared GPU memory;
- does not add shared GPU memory to system RAM when determining which model fits;
- caps automatic expert-cache sizing using currently available host memory;
- leaves 4 GiB for the OS and request-time CPU work.

Discrete AMD cards retain their existing independent-VRAM behavior.

## Implementation

- Add `gfx1151` to the HIP architecture allowlist and documentation.
- Enable the RDNA 3.5 signed dot-product path through `__builtin_amdgcn_sudot4`.
- Preserve signed zero in the Q2_0 half-dequantization path on gfx1151. The tested HIP compiler otherwise folds a negative scale multiplied by code zero to positive zero.
- Treat gfx1151 as the supported integrated-GPU exception during setup.
- Read the large KFD GTT pool instead of the small firmware VRAM carveout.
- Make model selection aware that gfx1151 GPU memory is shared with system RAM.
- Bound automatic expert-cache allocation using Linux host-memory availability.
- Use AMD’s current multi-architecture TheRock wheel index and select the required `device-gfx*` extras.
- Validate a candidate system ROCm installation with `rocminfo` and fall back to the wheels if its HSA runtime fails or crashes.
- Update the AMD setup tests and user documentation.

## Test system

- AMD Ryzen AI Max+ 395, 16 cores / 32 threads;
- AMD Radeon 8060S;
- `gfx1151`, wave32, 40 compute units;
- 128 GB unified memory;
- 116.0 GiB GPU-addressable shared GTT pool;
- CachyOS Linux with the kernel amdgpu driver.

## Tested HIP SDK

The successful build used the latest AMD nightly available at the time of validation on 2026-10-04:

[https://nightly.repo.amd.com/rocm/whl-next/](https://nightly.repo.amd.com/rocm/whl-next/)

Installed components:

```text
rocm[libraries,devel,device-gfx1151]
```

The SDK reports:

```text
$ bin/hipconfig
HIP version: 7.17.26392-0000000
```

`rocminfo` reports:

```text
ROCk module is loaded
=====================
HSA System Attributes
=====================
Runtime Version:         1.21
Runtime Ext Version:     1.32
System Timestamp Freq.:  1000.000000MHz
Sig. Max Wait Duration:  18446744073709551615 (0xFFFFFFFFFFFFFFFF)
Machine Model:           LARGE
System Endianness:       LITTLE
Mwaitx:                  ENABLED
XNACK enabled:           NO
DMAbuf Support:          YES
VMM Support:             YES
Fabric Support:          NO
```

The detected GPU agent is:

```text
Name:                    gfx1151
Marketing Name:          AMD Radeon 8060S Graphics
Device Type:             GPU
Memory Properties:       APU
Wavefront Size:          32
Compute Unit:            40
Max Clock Freq. (MHz):   2900
```

The build environment was rooted in the nightly wheel’s `_rocm_sdk_devel` directory. Environment variables pointing to another `/opt/rocm` installation were removed so that installation could not supply compiler or runtime libraries.

## Build and test validation

Using the nightly-wheel SDK:

- the complete source tree compiled for `gfx1151`;
- `strata-device` detected and exercised the GPU;
- native IQ expert kernels compiled;
- HIP MMQ compiled;
- 42 of 47 registered tests passed;
- four tests skipped normally;
- `ple_parity` was the only failure because its external model fixture was unavailable;
- no GPU test failed.

The setup-specific test suite also passes:

```text
$ python tools/test_setup_amd.py
Ran 30 tests in 0.017s

OK
```

An earlier validation using the working ROCm 7.13.99004 TheRock tree in `/opt/rocm-7.13` produced:

- complete engine and test-target compilation for `gfx1151`;
- passing `strata-device --list-devices`;
- passing `strata-device --selftest`;
- 58 of 65 registered tests passed;
- four tests skipped normally;
- the remaining three tests required unavailable model fixtures:
  `ple_parity`, `expert_parity`, and `pool_test`;
- no GPU test failed.

The passing GPU coverage included:

- plain hipBLAS prefill;
- MMQ prefill;
- native `v_dot4_i32_iu8`;
- mapped host memory;
- expert-cache staging;
- asynchronous handoff;
- Q2_0 signed-zero parity.

## End-to-end model validation

### Qwen3.8-Flash-Next Q2_0

Configuration:

- original GSQ-RCO Q2_0 model;
- context: 131,072 tokens;
- KV format: int8;
- 32,768 resident KV cells;
- MTP: `--spec 4 --spec-min-p 0.5`;
- one request at a time;
- CPU vision encoder available.

Startup:

- loaded the 31.64 GiB expert arena at 3.63 GiB/s;
- cached all 24,576 routed experts in the shared GPU pool;
- served OpenAI-compatible requests successfully.

Observed output performance:

- short responses: approximately 41.9–62.0 tok/s;
- reasoning-enabled response: 1,921 tokens in 41 seconds;
- sustained rate: 47.8 tok/s;
- live throughput remained around 47.6–48.9 tok/s through most of the long response.

### Unsloth UD-IQ4_XS

A second live configuration also passed:

- context: 65,536 tokens;
- KV format: int8;
- 32,768 resident KV cells;
- MTP enabled;
- all 24,576 routed experts cached;
- 55.43 GiB expert cache;
- CPU vision projector enabled.

A 2,205-token response completed at 36.1 tok/s. Other sustained responses measured approximately 34.6–36.1 tok/s.

Both Q2_0 and UD-IQ4_XS successfully processed image requests through the CPU vision encoder.

These are live validation measurements, not a standardized hardware benchmark. A separate community benchmark PR will follow after the controlled fresh-prompt and recall suites have been run.

## Known limitations

- Linux only. The ready-made Windows HIP archive does not include gfx1151.
- AMD GPU vision is not enabled; `--vision cpu` works.
- There is no gfx1151 hipBLASLt tuning table yet, so dense prompt projections use plain hipBLAS.
- Validation and performance observations are from one Radeon 8060S system.
- Shared-memory accounting is Linux-specific because it uses Linux host-memory and KFD topology information.

## Broken 7.14 nightly

The specific TheRock `7.14.0a20260612` nightly tested here is unusable on gfx1151.

The complete Strata tree compiles with it, but both its own `rocminfo` and `strata-device` segfault in:

```text
rocr::AMD::GpuAgent::InitDma()
```

during `hsa_init`, before Strata launches a GPU kernel.

Replacing only that installation’s `libhsa-runtime64.so.1` with the working ROCm 7.13 library makes `strata-device --selftest` and the same non-fixture GPU tests pass. This isolates the observed failure to that nightly’s HSA runtime. Mixing runtime components is useful for diagnosis only and is not a recommended installation.

Related upstream reports:

- [ROCm/TheRock#5763](https://github.com/ROCm/TheRock/issues/5763)
- [ROCm/TheRock#5779](https://github.com/ROCm/TheRock/issues/5779)

## AI assistance disclosure

Implementation review, documentation preparation, test interpretation, and commit organization were assisted by AI. All hardware builds, GPU tests, model runs, and reported measurements were executed on the Radeon 8060S system described above.

En el sitio

Enlaces a install, modelos, releases.