Issues / #1018

#1018 [gfx906] gr_up_fast_kernel<GrMulti> aborts the engine with HSA_STATUS_ERROR_MEMORY_APERTURE_VIOLATION on long-context chats (intermittent; the non-fast path is stable)

open · @drumblund · 0 comments · View on GitHub

BenchmarksSetup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants

Description

# [gfx906] `gr_up_fast_kernel<GrMulti>` aborts the engine with `HSA_STATUS_ERROR_MEMORY_APERTURE_VIOLATION` on long-context chats (intermittent; the non-fast path is stable)

## Environment

- GPU: 2× Radeon VII (gfx906), 16 GB each, `--layer-split 25`
- CPU/RAM: TR 2950X, 62 GB
- OS: Ubuntu 22.04, kernel 6.8.0-138-generic
- ROCm: `mixa3607/rocm-gfx906:7.14-complete` container (ROCm 7.14 community build with gfx906 kernels)
- Engine: 0.1.39, built from `6f32ec0` with `-DSTRATA_HIP_GFX906=ON -DCMAKE_HIP_ARCHITECTURES=gfx906`, plus two small compile fixes for gfx906 (available on request)
- Server args (abridged): `--mmap-experts --expert-cache auto --prefill 4096 --spec 4 --spec-min-p 0.5 --mtp ... --max-context 262144 --kv int8 --pool-workers 15 --layer-split 25 --trim-stage-weights --vram-reserve-mib 1024`
- Extra env in use: `STRATA_ARENA_MMAP=1`, `STRATA_NO_LARGEPAGES=1`, `kernel.numa_balancing=0`

## Symptom

The engine aborts (exit `-6`) and the server restarts it. Both occurrences named the same kernel:

```
Queue error: HSA_STATUS_ERROR_MEMORY_APERTURE_VIOLATION: The agent attempted to access memory beyond the largest legal address.
```

```
:0:rocdevice.cpp :4212: 56027459666 us:  Callback: Queue 0x70b938800000 aborting with error :
HSA_STATUS_ERROR_MEMORY_APERTURE_VIOLATION: The agent attempted to access memory beyond the largest legal address.
code: 0x29 [host: <host>, GPU index: 0,
kernel: strata::kernels::(anonymous namespace)::gr_up_fast_kernel(strata::kernels::(anonymous namespace)::GrMulti)]
```

The coredump block that follows:

```
GPU coredump: Pipe handler '/usr/share/apport/apport' not found or not executable.
Falling back to file-based dump. Set HSA_COREDUMP_PATTERN to override.
VGPU=0x5b28ea058670 SWq=0x70b93ab9e000, HWq=0x70b938800000, id=1
	Dispatch Header =0xb02 (type=2, barrier=1, acquire=1, release=1), setup=3
	grid=[40960, 1, 1], workgroup=[256, 1, 1]
	private_seg_size=0, group_seg_size=12288
	kernel_obj=0x70b933f86340, kernarg_address=0x0x70b715295500
	completion_signal=0x0, correlation_id=0
```

## Trigger conditions

- **Long-context chat**: prompts of ~190K–200K tokens (with prompt-cache reuse, so each turn re-reads only a few hundred tokens) and a few hundred generated tokens per turn.
- **Speculative decoding on** (`--spec 4`, MTP drafts).
- Intermittent: most turns are fine; two aborts happened 16 minutes apart after ~1 h of this workload.
- Reproduced twice in a row on 2026-10-06 (10:48, 11:04 local); the same kernel both times.

## Not address-space exhaustion

`dmesg` at boot says:

```
[drm] vm size is 262144 GB, 4 levels, block-size 9-bit, fragment size 9-bit
```

so the per-VM GPU VA aperture is 256 TB and the process's total GPU allocations are ~100 GB. This looks like a wild pointer / lifetime problem in the fast path rather than an aperture limit.

## Workaround

`STRATA_GR_FAST=0` (the non-fast `gr_up_multi_kernel` path) has been stable for hours under the same workload. Cost measured on a 4K-context, 300-token probe: **decode ~40 → ~32.5 tok/s (-18%)**, which is why I would prefer a real fix.

## Notes

- The fast path is gfx906-only (`#if defined(STRATA_HIP_GFX906)` in `src/kernels/cuda/fused_gr.cu`), so this will only show up on gfx906 builds.
- `gr_up_fast_kernel` (fused_gr.cu:904) is launched from `fused_gr_read_multi` (fused_gr.cu:814 and :1145) with `UPM_BLOCKS` blocks (160 for this model). The dispatch header in the dump shows `grid=[40960, 1, 1]`, which does not match `UPM_BLOCKS` — possibly the dump describes a neighbouring dispatch rather than the faulting one; flagging in case it helps.
- A second, unrelated gfx906 issue we hit is a checkpoint copy that races a stream capture (`operation would make the legacy stream depend on a capturing blocking stream`, worked around with `--prompt-cache-every 0`); happy to file it separately if useful.

Any pointers on where the fast path can compute an out-of-range address would be much appreciated — we can rebuild and test patches quickly.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.