Issues / #1407
#1407 [gfx1200, 2x RX 9060 XT layer split] prompt-read stalls (watchdog #29 aborts) + run-to-run prefill degradation: fresh prompts slow 3-4x mid-run, file tier re-reads 24-35 GB per request
open · @Iokkeh2025 · 0 comentários · No GitHub
BenchmarksServer & APIMulti-GPUAMD / HIPModels & quantsWindows
Descrição
[gfx1200, 2x RX 9060 XT layer split] prompt-read stalls (watchdog #29 aborts) + run-to-run prefill degradation: fresh prompts slow 3-4x mid-run, file tier re-reads 24-35 GB per request Environment | | | |-----------|---------------------------------------------------------------------------------------------------------------------------------| | Engine | v0.1.40.2, self-built, backend: hip, archs: [gfx1200] | | ROCm | 10.1.0-3 (gfx1200), HIP 7.16.26385, clang 24.0.0git | | GPUs | 2x RX 9060 XT 16 GB, --layer-split auto (0-23 / 24-47) | | Host | Ryzen 5 4600G (6C/12T), 32 GB DDR4, no swap, kernel 7.0.0-38 | | Model | Qwen3.8-Flash-Next IQ2_XS, ~39 GB experts, --mmap-experts | | hipBLASLt | self-tuned gfx1200-hipblaslt-100401.txt, log confirms tuning enabled (26 rows, version 100401) — table pairing is not the issue | Flags: --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp <rt> --max-context 65536 --kv q4_0 --mmap-experts --vram-reserve-mib 1024 --layer-split auto Env: STRATA_HIPBLASLT_TUNING, STRATA_GR_V3=1, later STRATA_IO_PREFETCH=1 STRATA_IO_STATS=1 1. Prompt-read stalls (#29 watchdog abort) Fresh 4K-38K prompts: the progress line freezes mid-chunk, the watchdog kills the engine: [strata] reading the prompt: 16,384 of 38,830 tokens, 495 s so far <- unchanged for 100+ s strata serve: no progress for 60 s during a request (reading the prompt (batched): layer 34 of the prompt chunk from token 0) ... (issue #29) -> exit -6 While frozen: both GPUs at 100% busy, 3 threads in D state at lock_mm_and_find_vma, 2 at kfd_wait_on_events, MemAvailable ~17 GB (no OOM). One run also hit verify: timed out at layer 8 (#267); an earlier pre-0.1.40.2 run wedged fully (empty HTTP reply, needed a kill). Freezes only ever occurred during the prompt pass, never decode. With STRATA_IO_PREFETCH=1 a full bench completed with no stalls (one data point — the prefetch path may be bypassing the fault storm). 2. File tier re-reads tens of GB per request (STRATA_IO_STATS) file tier I/O this request: the OS read 34,889.1 MB from storage (255,655 major faults) for 89,694.7 MB of expert reads file tier I/O this request: the OS read 24,306.9 MB ... page cache held 0.0 MB at the read io prefetch: 917 read ahead (1329.3 MB by pread), 0 used, 0 unused page cache held 0.0 MB even though ~17-20 GB is reclaimable — layer evictions (FADV_DONTNEED) make every request cold; 562 GB cumulative read for ~40 requests. Also 0 used / 0 unused for read-ahead looks like consumed preads aren't accounted for. 3. Prefill degrades within a run (same binary, same config) bench_prefill.py, back-to-back, engine otherwise idle: | trial | prompt | prefill | decode | |-------|--------|------------|--------| | 1 | 4,210 | 75.5 tok/s | 16.9 | | 2 | 8,830 | 87.5 tok/s | 21.0 | | 3 | 4,210 | 22.2 | 5.5 | | 4 | 8,830 | 33.5 | 34.9 | Directional, not random: first two prompts fast, last two 3-4x slower. Follow-ups inherit it — 114-tok follow-up after a 4,210 fresh = 56-59 tok/s, after an 8,830 fresh = 15-17 tok/s. A large fresh prompt seems to leave the file tier re-streaming what it just evicted. Ruled out OOM (none in journalctl), NVMe errors (SMART pass, 1.5 GB/s reads), GPU Xid/HSA errors (none in dmesg), VRAM exhaustion (~890 MiB free at load), hipBLASLt table mismatch (confirmed loaded each start). Notes - ReBAR aperture is 256 MB (board DSDT zeroes the windows; local issue — mentioned only because the mmap fault path is what a small aperture penalizes) - card0's memory clock sits at 96 MHz while DPM reads high mid-request (card1 stays at 1258 MHz); it does climb under full load Questions 1. When the prompt pass stalls at lock_mm_and_find_vma, is the watchdog catching the fault storm itself or a real hang? 2. Is page cache held 0.0 MB expected with --mmap-experts + FADV_DONTNEED evictions? 3. Are consumed prefetch preads miscounted (0 used, 0 unused)? Happy to run A/B cells o
No site
Links install, modelos, releases.