Issues / #1267

#1267 gfx1151 (Strix Halo) gets the ROCm 7.10.0a20251120 wheels, not the 7.14.1 that docs/STRIX_HALO.md was measured on: the engine crashes on kernel 7.2.8, and prompts run at half speed

closed · @alexdns1 · 4 comments · View on GitHub

BenchmarksSetup & installServer & APIAMD / HIPModels & quantsDocumentationLinux

Description

### What happens

`setup.py` installs one ROCm for every AMD card: `ROCM_VERSION = "7.10.0a20251120"` (setup.py:1493, added in
10d5cad for the RX 7900). Strix Halo support (c4a329f) added the gfx1151 wheel index but kept that version. Meanwhile
[docs/STRIX_HALO.md](docs/STRIX_HALO.md) says the gfx1151 work was built and measured on ROCm 7.14.1 (AMD's tarball),
and the shipped table `tools/hip/gfx1151-hipblaslt-100401.txt` is for 7.14.1's hipBLASLt (1.4.1).

On a Strix Halo with kernel 7.2.8 this has three effects:

1. **The server does not start.** The engine segfaults in the 7.10 wheel's HSA runtime on its first copy to the GPU,
   and the server only says "the engine exited before it was ready". The engine log has no error line (the crash
   is in the kernel log: `segfault ... in libhsa-runtime64.so.1`).
2. **No hipBLASLt table.** The 7.10 wheels carry hipBLASLt 1.2.0 (`100200`); there is no gfx1151 table for it, so the
   prompt's dense matrix products use plain hipBLAS.
3. **Setup's own test does not catch it.** `tools/test_setup_strix_halo.py` runs on recorded sysfs data, without a
   GPU, so nothing runs the wheels setup installs on a real gfx1151.

### Machine

Ryzen AI Max+ 395 (gfx1151, PCI 1002:1586), 128 GB, Fedora 44, kernel 7.2.8-200.fc44, linux-firmware 20260916,
boot parameters `amd_iommu=off amdgpu.gttsize=126976 ttm.pages_limit=32505856` (124 GiB GTT). Strata 82f46a8
(engine 0.1.40). Model: Unsloth UD-IQ4_XS, installed with `./setup.sh` (262144 context, int8 KV,
`--kv-resident 32768`, `--resident-budget-gib 55`).

### The crash

```
Thread 1 "strata" received signal SIGSEGV, Segmentation fault.
#0  rocr::AMD::GpuAgent::ReleaseQueueMainScratch(rocr::AMD::ScratchCache::ScratchInfo&) () from .venv/.../_rocm_sdk_devel/lib/libhsa-runtime64.so.1
#1  rocr::AMD::GpuAgent::QueueCreate(...) ()
#2  rocr::AMD::GpuAgent::InitDma()::$_0 ...
#5  rocr::HSA::hsa_queue_create(...) ()
#6-13 libamdhip64.so.7
#14 (anonymous namespace)::probe_pcie_h2d_gbps(...) ()
#15 main ()
```

`strace` shows the kernel refusing the runtime's copy queue just before it:
`ioctl(4, AMDKFD_IOC_CREATE_QUEUE, ...) = -1 EINVAL`, then `SIGSEGV ... si_addr=0x34`. It is not Strata's code: a
7-line Python `ctypes` program (`hipMalloc`, `hipHostMalloc`, `hipMemcpyAsync`) crashes the same way with the 7.10
wheel's `libamdhip64.so.7`.

The same program on other ROCm builds, same machine and kernel:

| ROCm | Result |
|---|---|
| 7.10.0a20251120 (setup's pin) | crash (above) |
| 7.11.0a20260121, 7.12.0a20260311, 7.13.0a20260515 nightlies | works |
| 7.14.0a20260529 to 7.14.0a20260608 nightlies | works |
| 7.14.0a20260609 to 7.14.0a20260612 nightlies | crash, a different one: `GpuAgent::InitDma` jumps to `0x100000001` during `hsa_init` |
| 7.14.1 release tarball (docs/STRIX_HALO.md) | works |
| Fedora's ROCm 7.1.1 (`rocm-runtime-7.1.1-6.fc44`) | works |

(Only the 7.10 and 7.14 crashes were traced with gdb. Neither is reported to AMD yet.)

### Speed

The engine setup compiled can run when its `lib_dirs` is emptied (it then loads Fedora's ROCm 7.1.1). Against an
engine built with docs/STRIX_HALO.md's own commands on ROCm 7.14.1 with the gfx1151 table, on the machine above:
8,192-token prompt, 256 output tokens, the run config's arguments with `--max-context 16384` and without
`--kv-resident`, one fresh `strata` process per run, the two engines interleaved, medians of 3.

| Engine | Prompt tok/s | Output tok/s |
|---|---|---|
| setup's (7.10 wheels build, run on ROCm 7.1.1, no table) | 262.9 | 52.24 |
| docs/STRIX_HALO.md build (ROCm 7.14.1, `gfx1151-hipblaslt-100401.txt`) | 547.9 | 54.40 |

Both printed the same 256 token ids in every run. Most of the prompt gain is plausibly the table, but the two
were not separated.

### What could fix it

- Install ROCm 7.14.1 for gfx1151 (the tarball docs/STRIX_HALO.md uses; sha256 in the doc), so that setup's engine,
  libraries and table match what the doc measured. AMD's release wheel index
  (`https://repo.amd.com/rocm/whl/gfx1151/`) stops at 7.13.0 today, so 7.14.1 means the tarball.
- Or pin gfx1151 to the 7.13.0 release wheels from that index. Not measured here, and there is no gfx1151 table for
  its hipBLASLt (unchecked which version it carries).
- Either way: a version per GPU family instead of one `ROCM_VERSION`, and a start check that tells the user when
  the engine dies in the ROCm runtime (the kernel log line) instead of only "exited before it was ready".

Local workaround (what this machine runs now): the engine built per docs/STRIX_HALO.md against ROCm 7.14.1 in
`~/rocm/install`, `lib_dirs` and `STRATA_HIPBLASLT_TUNING` in the run config pointing at it, and `engine/BUILD.json`
stamped with setup's source hash so that setup keeps the engine. After an update that changes the engine's source,
setup compiles it against the 7.10 wheels again and leaves the run config pointing at 7.14.1, so the workaround
has to be redone.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.