Pull requests / #1749

#1749 Two HIP fixes for Strix Halo (gfx1151): the wide top-k kernel and the VMM entry points

open · @vlupilin · 0 Kommentare · Auf GitHub

BenchmarksAMD / HIPNVIDIA / CUDAModels & quants

Beschreibung

# PR: two HIP fixes for Strix Halo (gfx1151) — the wide top-k kernel and the VMM entry points

Both changes are "a CUDA-only path that is dead on AMD for no hardware reason".
No behaviour change on NVIDIA cards; both are measured on a Ryzen AI Max+ 395
(Radeon 8060S, gfx1151, ROCm 10.1) with Qwen3.8-Flash-Next IQ3_XXS, 512K
context (YaRN 2), and reproduce with the repo's own tools (`qsa_select_bench`).

## 1. topk: compile the wide kernel on HIP too

`block_topk_wide_kernel` is plain shared-memory / atomicAdd code (no sm_90
feature), but its definition and dispatch branch sat behind
`#if !defined(__HIPCC__)`. On AMD every context past the register kernel's
reach (~135K cells) fell back to the 4-pass ref radix. The ids are identical
by construction — `qsa_select_bench` states 256/256 at 256 queries x 65537
blocks and 1/1 at nq=1.

Measured (`qsa_select_bench`, gfx1151):
- nq=1, 131073 blocks (a decode window at 512K context): 2.24 -> **0.77 ms** (2.9x)
- nq=256, 65537 blocks (a prompt batch at 262K): 3.16 -> **1.43 ms** (2.2x)

End-to-end on that PC this (with the scorer change of a follow-up) took the
decode of a 512K-token prompt from ~24-26 up and flattened the prefill curve
(515K: 473 -> 634 tok/s). `STRATA_TOPK_STREAM=0` still restores the register
kernel.

## 2. vmm: resolve the driver entry points under their hip- names on HIP

`vmm.cpp` looks up `cuMemAddressReserve` & co. by name; `libamdhip64` exports
the same entry points only as `hipMemAddressReserve`, `hipMemCreate`, ... — so
`vmm_available()` was always false on AMD and `--kv-grow` (and the VMM expert
cache) could never engage; the flag just printed "is off". The signatures
match, so on `__HIPCC__` the resolver now dlsym's the hip- name of the same
entry point.

With this, the APU loads `--kv-grow` with `--expert-cache 24576` (all experts,
49.8 GiB of GTT) and serves requests. Caveat we hit during bring-up: after one
K/V growth transition the first request hung once (watchdog restart); we run
this combination only under supervision, so the elastic paths deserve their
own validation pass on other cards before the default changes.

## Validation on our side

- `native_grouped_parity`, `hip_native_qsa_score`, `hip_prefill_fused_iq`:
  all pass (the top-k ids gate is inside `qsa_select_bench`).
- Needle retrieval 5/5 at 131K and 507K tokens after both changes.
- The engine's own `ExpertCache::verify_slot` byte-check passes.

Mehr auf der Site

Links zu Install, Modellen, Releases.