Issues / #1712
#1712 Strata on 2x RX 7900 XTX (HIP, AVX-only CPU): Memory access fault by GPU node-2 ... Page not present with concurrent requests Dates below are day/month (2026), times are local (CEST, UTC+2).
open · @mantovaniluca91 · 1 comentários · No GitHub
BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants
Descrição
## Summary
With **2 or more concurrent requests** (`"parallel": N`, layer split on two RX 7900 XTX) the engine dies with
`Memory access fault by GPU node-2 (Agent handle: ...) on address 0x7e74...000. Reason: Page not present or supervisor
privilege.` (exit code -6). The server restarts it by itself in 1-2 minutes. The faulting address is always a **host
(system RAM) virtual address**, and the fault always happens **when a conversation enters or leaves a batch slot**: right
after `strata batch: slot N takes ... tokens (copied in ... ms)`, `conversation cache: restored N tokens (checkpoint)` or
`strata batch: slot N gave back N tokens of this conversation (its turn checkpoint)`.
With 4 concurrent sessions it happens every 6-29 minutes. **It still happens on 0.1.41, with the copy engine off
(`HSA_ENABLE_SDMA=0`), with `STRATA_HIP_ADAPT_KERNEL_COPY=1`, and with the KV cache fully in VRAM (no `--kv-resident`).**
With one request at a time it was ~0.5/hour on the old two-socket placement, and 0 in ~3.5 h after moving both GPUs under
the same CPU.
**Side effect: the fault ends in a hardware reset of the card.** Right after the page fault the card's MES firmware stops
answering (`MES failed to response msg=3`, `failed to remove hardware queue from MES`), and amdgpu does a **MODE1 reset**
of that XTX (`VRAM is lost due to GPU reset!`; 9 resets of that card in ~22 h). The reset is then reported to **every ROCm
process on the machine**: four unrelated processes that only use the two MI50s (llama.cpp servers and a PyTorch service,
in other containers that cannot see the XTX) print `HW Exception by GPU node-4 (Agent handle: ...) reason :GPU Hang` at
the same second and exit (status 139 / SIGABRT). So one fault in Strata takes down every GPU service on the box for 1-2
minutes. The MES hang and the reset fan-out are driver/runtime behavior, not Strata's, but they make the fault expensive.
Kernel log below.
## Environment
| | |
|---|---|
| Strata | 0.1.40 / 0.1.40.1 (tests 06-09/10); 0.1.41 (fb58e0d) from 09/10 10:34 |
| Build | `local-hip`, archs gfx1100, `isa_floor: avx` (EXPERIMENTAL older-CPU build), vision cpu |
| GPUs used | 2x Radeon RX 7900 XTX 24 GB (gfx1100), layer split 27 (0-26 / 27-47), `--trim-stage-weights` |
| Other GPUs in the box | 2x Instinct MI50 32 GB (gfx906) running llama.cpp (ROCm 6.1.2), not visible to Strata's container |
| PCIe | Intel S2600CP (C602), PCIe 3.0 x8 slots; each XTX sits behind its own Navi switch |
| CPU / RAM | 2x Xeon E5-2630 v2 (Ivy Bridge, AVX, no AVX2), 12 cores, 2 NUMA nodes, 125 GB DDR3 |
| OS / kernel | Ubuntu 22.04.5, kernel 6.8.0-138-generic, in-kernel amdgpu, `iommu=pt`, `amdgpu.ppfeaturemask=0xffffffff` |
| ROCm | container `rocm/dev-ubuntu-24.04:7.2.4-complete` (engine built and run inside it, `--device /dev/kfd` + 2 render nodes) |
| Model | Qwen3.8-Flash-Next GSQ-RCO IQ3_XXS (ISTA-DASLab), native pack `iq3_xxs`, MTP rt, `--kv int8 --kv-resident 32768` |
| Server config | `reasoning_budget_tokens 4096`, `--max-context` 65536 / 131072 / 262144, `parallel` 1-4, `--batch-groups` 1/2/3/auto, optionally `--conversation-cache-mib 16384 --conversation-cache-slots N` |
| THP / NUMA | THP madvise; `kernel.numa_balancing` 1 (also tested with 0) |
## What we measured
| date | config | concurrent | duration | faults |
|---|---|---:|---|---:|
| 06-07/10 | 0.1.39-0.1.40, 64k, parallel 1, XTX on **two sockets** (node 0 + node 1) | 1 | ~15 h | 7 (~0.5/h), always at the start of a request after a pause |
| 08/10 day | 0.1.40, XTX moved under the **same socket**, 64k-262k, parallel 1/2/4, batch-groups 1-2, no parking | 1-4, short tests | ~3.5 h engine time | **0** |
| 08/10 17:38 | 262k, parallel 4, batch-groups 2, **parking 4x16 GiB** | 4 agent sessions | 75 min | **2** (17:58, 18:39), both right after `conversation cache: restored ... (checkpoint)` |
| 09/10 08:38 | same, **no parking** | 4 | ~35 min | **1** after ~15 min, right after `strata batch: slot 2 takes 21599 tokens` |
| 09/10 09:25 | same as 08/10 17:38 (parking) with **`kernel.numa_balancing=0`** | 4 | ~15 min | **1** after ~6 min, right after `conversation cache: parked/restored` |
| 09/10 09:52 | 262k, parallel 3, batch-groups 3, parking 3 | 3 | 22 min | **1** after 22 min |
| 09/10 10:34 | **0.1.41**, 262k, parallel 4, batch-groups auto, parking 4x16 GiB, without `--remote-expert-opt` | 4 | 29 min | **1** after 29 min, right after `conversation cache: restored 26338 tokens (checkpoint)` |
| 09/10 11:06 | 0.1.41, same, parallel 3 | 3 | 24 min | **1** after 24 min, right after `conversation cache: parked` |
| 09/10 11:37 | 0.1.41, 4 at 262k + `STRATA_HIP_ADAPT_KERNEL_COPY=1` (its startup line was in the log) | 4 | 25 min | **1** after 25 min, right after `conversation cache: restored 26552 tokens (checkpoint)` |
| 09/10 12:14 | 0.1.41, 4 at 262k + `HSA_ENABLE_SDMA=0` (checked in the engine's `/proc/<pid>/environ`) | 4 | 19 min | **1** after 19 min, right after `conversation cache: restored 30664 tokens (checkpoint)` |
| 09/10 12:42 | 0.1.41, 4 at 262k, **no `--kv-resident`** (KV fully in VRAM, 0 `KV streaming` lines; slot sessions 7.5 GiB per card) | 4 | 30 min | **0** (37 restores, 116 slot admissions) |
| 09/10 13:15 | 0.1.41, same (KV in VRAM) at **131k** | 4 | 27 min | **1** after 27 min, right after `strata batch: slot 2 gave back 25734 tokens of this conversation (its turn checkpoint)` |
Excluded so far: GPUs on different sockets (both XTX now on NUMA node 0), conversation parking (09/10 08:38), automatic
NUMA balancing, the SDMA copy engine (both switches), KV streaming, the 0.1.40 -> 0.1.41 update. With the KV in VRAM the
faults may be somewhat rarer (1 in 57 min of 4-session load, ~120 restores) but they do not stop.
Not yet tested: `HSA_USERPTR_FOR_PAGED_MEM=0`, `GPU_PINNED_MIN_XFER_SIZE=1048576`, `--mmap-experts`.
One instance per card (no layer split) could not be judged: it runs, but too slowly on this CPU (see below), so in 10
minutes there were too few slot admissions to say anything about the fault.
## Log excerpts
0.1.41, KV fully in VRAM (no `--kv-resident`), 131k, 4 sessions (09/10 13:42 local):
```
strata batch: slot 0 takes 31592 tokens (copied in 278.8 ms)
strata serve: prompt 31592 tokens = 31591 reused + 1 read in 6 ms (179.1 tok/s), 1 generated in 27 ms (37.5 tok/s), drafts accepted 0 of 0, 0 checkpoints
strata serve: decode expert cache hit rate: 100.0% (480 hits / 480 lookups)
strata serve: conversation cache: parked 31592 tokens in 200.5 ms; parked=4 bytes=4109642032 evictions=46 snapshot_bytes=600915304 reused_kv_bytes=0
strata batch: slot 2 gave back 25734 tokens of this conversation (its turn checkpoint) in 253.1 ms
Memory access fault by GPU node-2 (Agent handle: 0x94e5490) on address 0x732b48cce000. Reason: Page not present or supervisor privilege.
```
0.1.41, `HSA_ENABLE_SDMA=0`, 262k with KV streaming (09/10 12:33 local):
```
strata serve: conversation cache: dropped 1 superseded copy of this conversation; parked=2
strata serve: conversation cache: parked 56599 tokens in 787.4 ms; parked=3 bytes=3569736360 evictions=15 snapshot_bytes=1220564336 reused_kv_bytes=0
strata serve: conversation cache: restored 30664 tokens (checkpoint) in 230.0 ms; parked=3 bytes=4586586976
Memory access fault by GPU node-2 (Agent handle: 0x8b054a0) on address 0x7a39932e1000. Reason: Page not present or supervisor privilege.
```
0.1.40, numa_balancing=0 (09/10 09:34 local):
```
strata serve: conversation cache: restored 30012 tokens (checkpoint) in 192.5 ms; parked=4 bytes=7663099208
strata batch: slot 0 takes 31087 tokens (copied in 540.8 ms)
strata serve: prompt 31087 tokens = 30012 reused + 1075 read in 2469 ms (435.3 tok/s), 1 generated in 30 ms (33.3 tok/s), drafts accepted 0 of 0, 5 checkpoints
strata serve: decode expert cache hit rate: 100.0% (480 hits / 480 lookups)
strata serve: KV streaming: 54.89% of 15402 block reads hit VRAM, 28.0 MiB read from RAM
strata serve: conversation cache: parked 31087 tokens in 277.3 ms; parked=4 bytes=6721475160 evictions=19 snapshot_bytes=1622688976 reused_kv_bytes=458098944
Memory access fault by GPU node-2 (Agent handle: 0x1e5a3400) on address 0x7e747485b000. Reason: Page not present or supervisor privilege.
```
Without parking (09/10 08:57):
```
strata batch: slot 2 takes 21599 tokens (copied in 490.4 ms)
strata serve: prompt 21599 tokens = 17796 reused + 3803 read in 3382 ms (1124.5 tok/s), 1 generated in 28 ms (35.8 tok/s), drafts accepted 0 of 0, 4 checkpoints
strata serve: decode expert cache hit rate: 100.0% (480 hits / 480 lookups)
strata serve: KV streaming: 97.11% of 1716048 block reads hit VRAM, 199.5 MiB read from RAM
Memory access fault by GPU node-2 (Agent handle: 0xd5ee3c0) on address 0x70d707d16000. Reason: Page not present or supervisor privilege.
```
Server side, every time: `the engine stopped unexpectedly (exit code -6)` once per in-flight request, then
`the engine had stopped (exit code -6); starting it again`. No OOM: 60+ GB of RAM were available during the tests.
`node-2` is the first GPU visible in the container, i.e. the XTX running layers 0-26 (`node-4` / PCI `0000:09:00.0` on the host).
Kernel log of the same fault (09/10 13:43:58 local, 0.1.41, KV in VRAM, 131k; the last faulting address is the one the
engine reported, `0x732b48cce000`):
```
amdgpu 0000:09:00.0: amdgpu: [gfxhub] page fault (src_id:0 ring:24 vmid:8 pasid:32783, for process strata pid 2716527 thread strata pid 2716527)
amdgpu 0000:09:00.0: amdgpu: in page starting at address 0x0000732b48cce000 from client 10
amdgpu 0000:09:00.0: amdgpu: GCVM_L2_PROTECTION_FAULT_STATUS:0x00000000
amdgpu 0000:09:00.0: amdgpu: Faulty UTCL2 client ID: CB/DB (0x0)
amdgpu 0000:09:00.0: amdgpu: MORE_FAULTS: 0x0
amdgpu 0000:09:00.0: amdgpu: WALKER_ERROR: 0x0
amdgpu 0000:09:00.0: amdgpu: PERMISSION_FAULTS: 0x0
amdgpu 0000:09:00.0: amdgpu: MAPPING_ERROR: 0x0
amdgpu 0000:09:00.0: amdgpu: RW: 0x0
(five such faults in the same second, pages 0x732acdc6b000, 0x732ad92af000, 0x732ad57df000, 0x732b48cce000, 0x732ac96e4000)
[drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3
amdgpu 0000:09:00.0: amdgpu: failed to remove hardware queue from MES, doorbell=0x1004
amdgpu 0000:09:00.0: amdgpu: MES might be in unrecoverable state, issue a GPU reset
amdgpu 0000:09:00.0: amdgpu: Failed to evict queue 3 (then 2, 1, 0)
amdgpu 0000:09:00.0: amdgpu: GPU reset begin!
amdgpu: sq_intr: error, detail 0x00000000, type 2, sh 1, priv 1, wave_id 0, simd_id 0, wgp_id 0 (x10)
[drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue (x8)
amdgpu 0000:09:00.0: amdgpu: MODE1 reset
[drm] VRAM is lost due to GPU reset!
amdgpu 0000:09:00.0: amdgpu: GPU reset(9) succeeded!
```
Other ROCm processes (MI50 only), same second:
```
HW Exception by GPU node-4 (Agent handle: 0x573d53a658e0) reason :GPU Hang (x4, then exit status 139 / SIGABRT)
```
## Other findings that may help
1. **`/metrics` `totals.decode_ms` is far too small with batching.** 12 batched requests wrote 3,711 tokens, and the sum of
their `decode_ms` is 323.6 ms, while each request's `decode_tok_s` says 34-36 tok/s (~109 s in total).
`output_tokens / decode_ms` then gives "750 tok/s". The per-request `decode_tok_s` looks right.
2. **The GPU image encoder works on AMD with a stock build.** `tools/vision` built with `-DGGML_HIP=ON -DAMDGPU_TARGETS=gfx906
-DCMAKE_EXE_LINKER_FLAGS=-no-pie` (ROCm 6.1.2; `-no-pie` is needed because `libggml-hip.a` is not PIC), started by
`server.py` through a wrapper that sets `LD_LIBRARY_PATH` to ROCm 6.1, with `"vision": {"gpu": true, "cuda_device": 0}`
on a spare MI50 and `env.HIP_VISIBLE_DEVICES = "1,2"` for the engine.
- Speed: 1.1 s per image instead of 91 s on the CPU, +1.4 GiB of VRAM.
- Output: per-row cosine vs the CPU encoder, median 0.9993 (min 0.954).
- Real task, 16 delivery notes: 17.4 s per document instead of 150 s; 32/No site
Links install, modelos, releases.