Issues / #1407

#1407 [gfx1200, 2x RX 9060 XT layer split] prompt-read stalls (watchdog #29 aborts) + run-to-run prefill degradation: fresh prompts slow 3-4x mid-run, file tier re-reads 24-35 GB per request

open · @Iokkeh2025 · 0 comentarios · En GitHub

BenchmarksServer & APIMulti-GPUAMD / HIPModels & quantsWindows

Descripción

[gfx1200, 2x RX 9060 XT layer split] prompt-read stalls (watchdog #29 aborts) + run-to-run prefill degradation: fresh prompts slow 3-4x mid-run, file tier re-reads 24-35 GB per request

Environment

|           |                                                                                                                                 |
|-----------|---------------------------------------------------------------------------------------------------------------------------------|
| Engine    | v0.1.40.2, self-built, backend: hip, archs: [gfx1200]                                                                           |
| ROCm      | 10.1.0-3 (gfx1200), HIP 7.16.26385, clang 24.0.0git                                                                             |
| GPUs      | 2x RX 9060 XT 16 GB, --layer-split auto (0-23 / 24-47)                                                                          |
| Host      | Ryzen 5 4600G (6C/12T), 32 GB DDR4, no swap, kernel 7.0.0-38                                                                    |
| Model     | Qwen3.8-Flash-Next IQ2_XS, ~39 GB experts, --mmap-experts                                                                       |
| hipBLASLt | self-tuned gfx1200-hipblaslt-100401.txt, log confirms tuning enabled (26 rows, version 100401) — table pairing is not the issue |

Flags: --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp <rt> --max-context 65536 --kv q4_0 --mmap-experts --vram-reserve-mib 1024 --layer-split auto
Env: STRATA_HIPBLASLT_TUNING, STRATA_GR_V3=1, later STRATA_IO_PREFETCH=1 STRATA_IO_STATS=1

1. Prompt-read stalls (#29 watchdog abort)

Fresh 4K-38K prompts: the progress line freezes mid-chunk, the watchdog kills the engine:


[strata] reading the prompt: 16,384 of 38,830 tokens, 495 s so far   <- unchanged for 100+ s
strata serve: no progress for 60 s during a request (reading the prompt (batched):
  layer 34 of the prompt chunk from token 0) ... (issue #29)   -> exit -6


While frozen: both GPUs at 100% busy, 3 threads in D state at lock_mm_and_find_vma, 2 at kfd_wait_on_events, MemAvailable ~17 GB (no OOM). One run also hit verify: timed out at layer 8 (#267); an earlier pre-0.1.40.2 run wedged fully (empty HTTP reply, needed a kill). Freezes only ever occurred during the prompt pass, never decode.

With STRATA_IO_PREFETCH=1 a full bench completed with no stalls (one data point — the prefetch path may be bypassing the fault storm).

2. File tier re-reads tens of GB per request (STRATA_IO_STATS)


file tier I/O this request: the OS read 34,889.1 MB from storage (255,655 major faults)
  for 89,694.7 MB of expert reads
file tier I/O this request: the OS read 24,306.9 MB ... page cache held 0.0 MB at the read
  io prefetch: 917 read ahead (1329.3 MB by pread), 0 used, 0 unused


page cache held 0.0 MB even though ~17-20 GB is reclaimable — layer evictions (FADV_DONTNEED) make every request cold; 562 GB cumulative read for ~40 requests. Also 0 used / 0 unused for read-ahead looks like consumed preads aren't accounted for.

3. Prefill degrades within a run (same binary, same config)

bench_prefill.py, back-to-back, engine otherwise idle:

| trial | prompt | prefill    | decode |
|-------|--------|------------|--------|
| 1     | 4,210  | 75.5 tok/s | 16.9   |
| 2     | 8,830  | 87.5 tok/s | 21.0   |
| 3     | 4,210  | 22.2       | 5.5    |
| 4     | 8,830  | 33.5       | 34.9   |

Directional, not random: first two prompts fast, last two 3-4x slower. Follow-ups inherit it — 114-tok follow-up after a 4,210 fresh = 56-59 tok/s, after an 8,830 fresh = 15-17 tok/s. A large fresh prompt seems to leave the file tier re-streaming what it just evicted.

Ruled out

OOM (none in journalctl), NVMe errors (SMART pass, 1.5 GB/s reads), GPU Xid/HSA errors (none in dmesg), VRAM exhaustion (~890 MiB free at load), hipBLASLt table mismatch (confirmed loaded each start).

Notes

- ReBAR aperture is 256 MB (board DSDT zeroes the windows; local issue — mentioned only because the mmap fault path is what a small aperture penalizes)
- card0's memory clock sits at 96 MHz while DPM reads high mid-request (card1 stays at 1258 MHz); it does climb under full load

Questions

1. When the prompt pass stalls at lock_mm_and_find_vma, is the watchdog catching the fault storm itself or a real hang?
2. Is page cache held 0.0 MB expected with --mmap-experts + FADV_DONTNEED evictions?
3. Are consumed prefetch preads miscounted (0 used, 0 unused)?

Happy to run A/B cells o

En el sitio

Enlaces a install, modelos, releases.