Pull requests / #1749
#1749 Two HIP fixes for Strix Halo (gfx1151): the wide top-k kernel and the VMM entry points
open · @vlupilin · 0 comentarios · En GitHub
BenchmarksAMD / HIPNVIDIA / CUDAModels & quants
Descripción
# PR: two HIP fixes for Strix Halo (gfx1151) — the wide top-k kernel and the VMM entry points Both changes are "a CUDA-only path that is dead on AMD for no hardware reason". No behaviour change on NVIDIA cards; both are measured on a Ryzen AI Max+ 395 (Radeon 8060S, gfx1151, ROCm 10.1) with Qwen3.8-Flash-Next IQ3_XXS, 512K context (YaRN 2), and reproduce with the repo's own tools (`qsa_select_bench`). ## 1. topk: compile the wide kernel on HIP too `block_topk_wide_kernel` is plain shared-memory / atomicAdd code (no sm_90 feature), but its definition and dispatch branch sat behind `#if !defined(__HIPCC__)`. On AMD every context past the register kernel's reach (~135K cells) fell back to the 4-pass ref radix. The ids are identical by construction — `qsa_select_bench` states 256/256 at 256 queries x 65537 blocks and 1/1 at nq=1. Measured (`qsa_select_bench`, gfx1151): - nq=1, 131073 blocks (a decode window at 512K context): 2.24 -> **0.77 ms** (2.9x) - nq=256, 65537 blocks (a prompt batch at 262K): 3.16 -> **1.43 ms** (2.2x) End-to-end on that PC this (with the scorer change of a follow-up) took the decode of a 512K-token prompt from ~24-26 up and flattened the prefill curve (515K: 473 -> 634 tok/s). `STRATA_TOPK_STREAM=0` still restores the register kernel. ## 2. vmm: resolve the driver entry points under their hip- names on HIP `vmm.cpp` looks up `cuMemAddressReserve` & co. by name; `libamdhip64` exports the same entry points only as `hipMemAddressReserve`, `hipMemCreate`, ... — so `vmm_available()` was always false on AMD and `--kv-grow` (and the VMM expert cache) could never engage; the flag just printed "is off". The signatures match, so on `__HIPCC__` the resolver now dlsym's the hip- name of the same entry point. With this, the APU loads `--kv-grow` with `--expert-cache 24576` (all experts, 49.8 GiB of GTT) and serves requests. Caveat we hit during bring-up: after one K/V growth transition the first request hung once (watchdog restart); we run this combination only under supervision, so the elastic paths deserve their own validation pass on other cards before the default changes. ## Validation on our side - `native_grouped_parity`, `hip_native_qsa_score`, `hip_prefill_fused_iq`: all pass (the top-k ids gate is inside `qsa_select_bench`). - Needle retrieval 5/5 at 131K and 507K tokens after both changes. - The engine's own `ExpertCache::verify_slot` byte-check passes.
En el sitio
Enlaces a install, modelos, releases.