Issues / #1018
#1018 [gfx906] gr_up_fast_kernel<GrMulti> aborts the engine with HSA_STATUS_ERROR_MEMORY_APERTURE_VIOLATION on long-context chats (intermittent; the non-fast path is stable)
open · @drumblund · 0 commentaires · Sur GitHub
BenchmarksSetup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants
Description
# [gfx906] `gr_up_fast_kernel<GrMulti>` aborts the engine with `HSA_STATUS_ERROR_MEMORY_APERTURE_VIOLATION` on long-context chats (intermittent; the non-fast path is stable) ## Environment - GPU: 2× Radeon VII (gfx906), 16 GB each, `--layer-split 25` - CPU/RAM: TR 2950X, 62 GB - OS: Ubuntu 22.04, kernel 6.8.0-138-generic - ROCm: `mixa3607/rocm-gfx906:7.14-complete` container (ROCm 7.14 community build with gfx906 kernels) - Engine: 0.1.39, built from `6f32ec0` with `-DSTRATA_HIP_GFX906=ON -DCMAKE_HIP_ARCHITECTURES=gfx906`, plus two small compile fixes for gfx906 (available on request) - Server args (abridged): `--mmap-experts --expert-cache auto --prefill 4096 --spec 4 --spec-min-p 0.5 --mtp ... --max-context 262144 --kv int8 --pool-workers 15 --layer-split 25 --trim-stage-weights --vram-reserve-mib 1024` - Extra env in use: `STRATA_ARENA_MMAP=1`, `STRATA_NO_LARGEPAGES=1`, `kernel.numa_balancing=0` ## Symptom The engine aborts (exit `-6`) and the server restarts it. Both occurrences named the same kernel: ``` Queue error: HSA_STATUS_ERROR_MEMORY_APERTURE_VIOLATION: The agent attempted to access memory beyond the largest legal address. ``` ``` :0:rocdevice.cpp :4212: 56027459666 us: Callback: Queue 0x70b938800000 aborting with error : HSA_STATUS_ERROR_MEMORY_APERTURE_VIOLATION: The agent attempted to access memory beyond the largest legal address. code: 0x29 [host: <host>, GPU index: 0, kernel: strata::kernels::(anonymous namespace)::gr_up_fast_kernel(strata::kernels::(anonymous namespace)::GrMulti)] ``` The coredump block that follows: ``` GPU coredump: Pipe handler '/usr/share/apport/apport' not found or not executable. Falling back to file-based dump. Set HSA_COREDUMP_PATTERN to override. VGPU=0x5b28ea058670 SWq=0x70b93ab9e000, HWq=0x70b938800000, id=1 Dispatch Header =0xb02 (type=2, barrier=1, acquire=1, release=1), setup=3 grid=[40960, 1, 1], workgroup=[256, 1, 1] private_seg_size=0, group_seg_size=12288 kernel_obj=0x70b933f86340, kernarg_address=0x0x70b715295500 completion_signal=0x0, correlation_id=0 ``` ## Trigger conditions - **Long-context chat**: prompts of ~190K–200K tokens (with prompt-cache reuse, so each turn re-reads only a few hundred tokens) and a few hundred generated tokens per turn. - **Speculative decoding on** (`--spec 4`, MTP drafts). - Intermittent: most turns are fine; two aborts happened 16 minutes apart after ~1 h of this workload. - Reproduced twice in a row on 2026-10-06 (10:48, 11:04 local); the same kernel both times. ## Not address-space exhaustion `dmesg` at boot says: ``` [drm] vm size is 262144 GB, 4 levels, block-size 9-bit, fragment size 9-bit ``` so the per-VM GPU VA aperture is 256 TB and the process's total GPU allocations are ~100 GB. This looks like a wild pointer / lifetime problem in the fast path rather than an aperture limit. ## Workaround `STRATA_GR_FAST=0` (the non-fast `gr_up_multi_kernel` path) has been stable for hours under the same workload. Cost measured on a 4K-context, 300-token probe: **decode ~40 → ~32.5 tok/s (-18%)**, which is why I would prefer a real fix. ## Notes - The fast path is gfx906-only (`#if defined(STRATA_HIP_GFX906)` in `src/kernels/cuda/fused_gr.cu`), so this will only show up on gfx906 builds. - `gr_up_fast_kernel` (fused_gr.cu:904) is launched from `fused_gr_read_multi` (fused_gr.cu:814 and :1145) with `UPM_BLOCKS` blocks (160 for this model). The dispatch header in the dump shows `grid=[40960, 1, 1]`, which does not match `UPM_BLOCKS` — possibly the dump describes a neighbouring dispatch rather than the faulting one; flagging in case it helps. - A second, unrelated gfx906 issue we hit is a checkpoint copy that races a stream capture (`operation would make the legacy stream depend on a capturing blocking stream`, worked around with `--prompt-cache-every 0`); happy to file it separately if useful. Any pointers on where the fast path can compute an out-of-range address would be much appreciated — we can rebuild and test patches quickly.
Sur le site
Liens install, modèles, releases.