Issues / #1691
#1691 gfx906: 0.1.41 verify timeouts ("timed out at layer N"); prompt-read stalls are pre-existing (clamp blame retracted)
open · @KevinX8 · 1 コメント · GitHub で見る
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
本文
Bisected a gfx906 regression to the q8_1 clamp pair. Reporter setup below, A/B table, then logs. ## Setup - 1x AMD Instinct MI50 16 GB (gfx906, wave64), Ryzen 9 5900X (AVX2, no AVX-512), 60 GB RAM, Ubuntu 26.10, kernel amdgpu - ROCm 10.1 community gfx906 build (also reproduced on 10.0), HIP 7.16, clang 24 - Engine built with `-DSTRATA_HIP_GFX906=ON -DCMAKE_HIP_ARCHITECTURES=gfx906` (docs/AMD_HIP.md recipe) - Model: Unsloth UD-IQ4_XS (native pack), 262144 context, q4_0 KV + `--kv-resident 32768`, `--resident-budget-gib 43`, MTP `--spec 4 --spec-min-p 0.5`, single GPU - 0.1.40.x needs one local patch to build (two defines in the gfx906 compat header, same class as #1396): `cudaStreamCreateWithPriority` / `cudaDeviceGetStreamPriorityRange` in `include/strata/platform/hip_compat/strata_hip.h` (a732a19 fixed the neighbouring `cudaEventBlockingSync` but not these; `src/core/mtp.cpp:423-424` uses them) ## What 0.1.41 does (fb58e0d) Two failure signatures on the same box where 0.1.40.x serves all day: 1. `Unrecoverable native verification failure: verify: timed out at layer 10` and `... layer 26; its GPU waits were released` — decode stalls (18 -> 4.7 tok/s) then the engine exits, twice. 2. `no progress for 60 s during a request (reading the prompt (batched): waiting for the GPU (attention, router) ...` — twice more. Decode on the runs that complete: ~19.8 tok/s vs ~26-28 tok/s on 0.1.40.x. ## Bisect (v0.1.40.4 base, same flags/config/machine) - v0.1.40.4 alone: story x3 + a full live session, no hang, 23-28 tok/s decode. - v0.1.40.4 + ccd657a (HC norm scratch): 3 stories, 29.0/27.6 tok/s after warmup, no hang. Exonerated. - v0.1.40.4 + e372d4f + 7a0e302 (q8_1 clamps, isolated, no ccd657a): 4 consecutive prompt-read hangs across mixed traffic — a 27-token prompt dying 12 tokens in, a 2.6K prompt dying 5 tokens from the end, a 3.8K prompt dying 5 tokens from the end — every one `waiting for the GPU (attention, router)`, each ended by the 60 s watchdog + engine restart. Nothing else served in between except the hangs. So the clamp pair alone reproduces a hang on gfx906. The 0.1.41 verify-timeout crashes were not re-tested to 5x (stopped per the reporter's rule after the bisect landed); they may be the same root cause surfacing in decode. Happy to re-run anything on this card (single MI50 16 GB, can build any commit with the one-line patch above).
関連リンク
インストール・モデル・リリースへの站内リンク。