Issues / #937

#937 `[gfx1030] HSA_STATUS_ERROR_MEMORY_APERTURE_VIOLATION in gdn_step_commit_kernel after a few long fresh prefills`

open · @lucaravera-glitch · 2 Kommentare · Auf GitHub

BenchmarksSetup & installServer & APIAMD / HIPModels & quants

Beschreibung

## Summary

On an RX 6800 (gfx1030, Navi 21) the engine reliably dies after a small number of
long, *fresh* (non-KV-cached) prefills, with:

```
Queue error: HSA_STATUS_ERROR_MEMORY_APERTURE_VIOLATION: The agent attempted to access memory beyond the largest legal address.
```

The faulting kernel is always `strata::kernels::gdn_step_commit_kernel`, and the server
surfaces it as `the engine stopped unexpectedly (exit code -13)` followed by an automatic
engine restart.

Short/cached sessions are stable — a full 5/5 agent benchmark (~20 requests) runs clean.
The failure scales with **fresh prefill volume per engine lifetime**, not with
`--max-context`: prompts of 10k tokens fault just as readily as 34.5k, after 2-7 fresh
prefills.

## Environment

| | |
|---|---|
| Strata | `v0.1.39`, commit `6f32ec0`, engine built from source |
| GPU | Radeon RX 6800, 16 GB, `1002:73BF`, gfx1030 (Navi 21) |
| BAR0 | 16 GiB (`Region 0 ... [size=16G]`) — 256 MiB before enabling Above 4G Decoding |
| Driver | `amdgpu` from kernel `7.2.5-3-omarchy` (Omarchy/Arch) |
| Userspace | ROCm `7.14.0a20260612` (community gfx103X wheel), matching the kernel vintage |
| Model | Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S, 2 shards, 83.6 GB, 46.84 GiB expert arena |
| RAM | 94 GiB |

Launch args:

```
--max-context 65536 --kv int8 --kv-resident 65536 --vram-reserve-mib 3072
--prefill 2048 --pcie-frac 0 --adapt-every 0 --spec 4 --mtp <dir>
```

## Reproduction

1. Start the server and let the engine load.
2. Issue **fresh** prefills of ~10k-34.5k tokens (a probe that nonces both ends of the
   prompt guarantees no KV reuse).
3. Repeat in a loop.

Expect a fault within 2-7 completed prefills. Prefill throughput before the fault is
normal (~330 tok/s); nothing degrades, it simply dies between requests.

Representative run (34,507-token fresh prefills, faulting lifetime marked):

```
engine life #1: [1:34507tok 34507 fresh] [2:34507tok fresh] [3:34507tok fresh] [4:34507tok fresh]  <<FAULT>>
engine life #2: [1:34507tok fresh] [2:34507tok fresh]                                            <<FAULT>>
engine life #3: [1:34507tok fresh] [2..7:10063tok fresh]                                       <<FAULT>>
engine life #4: [1:10063tok fresh] [2:10063tok fresh] [3:10063tok fresh]
```

Aggregate: **16 clean / 3 faults across 4 engine lifetimes.**

## Ruled out

| Hypothesis | Test | Result |
|---|---|---|
| VRAM exhaustion | `--vram-reserve-mib 3072` (2990 MiB free at load) | no change |
| Prefill chunk-boundary bug | `--prefill 2048` | fixed a *different*, reproducible hang at the 8193 boundary; aperture fault unaffected |
| Small BAR (`Region 0 = 256M`) | enabled Above 4G Decoding → 16 GiB BAR | no change |
| Expert-arena mapping | `vm.nr_hugepages=24000`, arena switches to `hugetlb 2 MB pages` | no change |
| KV streaming / VRAM paging | `--kv-resident 65536`, no streaming at startup | no change |
| Host RAM exhaustion | 35-50 GiB available throughout; `free` never near OOM | no change |
| ROCm/driver vintage skew | kernel and ROCm are contemporary; bare metal | no change |

## Contrast case

With prompts that ride the KV cache (only the first turn prefills ~8k, later turns reuse),
a full agent benchmark ran **5/5 tasks, ~20 requests, 0 faults**:

```
strata serve: prompt 8090 tokens = 32768 reused + 0 read ...
```

So the practical risk scales with how often fresh long contexts are started, not with
session length or `--max-context`.

## Related: confirms #884

On this same setup, before settling on `--pcie-frac 0 --adapt-every 0`, prefill stalled
indefinitely (`said nothing for 600 s`) in roughly 1 in 3 long requests — the
"verify hangs after a long prompt" signature in #884. Two of the three knobs named there
(adaptive expert swaps, PCIe sharing) being off does remove that hang, which is offered as
confirmation data for #884. The aperture fault reported here survives both flags, so it
looks like a separate defect.

## Workarounds so far

- Cached/short sessions: reliable (5/5 benchmark, 0 faults).
- The server auto-restarts the engine, so the next request succeeds, but the ~160 s
  reload is disruptive mid-session.
- No configuration of `--max-context`, BAR size, or hugepages removes the fault.

Happy to provide full logs (`strata-iq3_s.log`, `strata-server.out`) and a standalone
repro script if useful.

Mehr auf der Site

Links zu Install, Modellen, Releases.