Issues / #1500
#1500 Engine stopped unexpectedly on strix halo
open · @stiff · 0 コメント · GitHub で見る
BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPModels & quantsSecurityWindows
本文
Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S 128k context, after doing ~50 turns session: ``` [strata] reading the prompt: 95,573 of 95,578 tokens, 112 s so far [strata] reading the prompt: 95,573 of 95,578 tokens, 122 s so far [strata] reading the prompt: 95,573 of 95,578 tokens, 132 s so far [strata] reading the prompt: 95,573 of 95,578 tokens, 142 s so far [strata] the engine stopped unexpectedly (exit code -6). The engine stopped itself because it had stopped making progress - a hang it caught. Its log line: strata serve: no progress for 60 s during a request (reading the prompt (batched), done up to token 95573) - stopping the engine so the server starts it again (issue #29) - please report it at github.com/Niko1221/Strata/issues. The next request starts the engine again. Its log: /home/stiff/tmp/strata/strata-iq3_s.log ``` ROCm 7.14 from AMD Debian repo, log snippet: ``` strata: gfx1151 (Strix Halo): 18 exact speed switches on by default (STRATA_GFX1151_DEFAULTS=0 turns them off; a switch you set is kept): GDN_HEAD GDN_PP GDN_CONVL2 GDN_NOY CVEC_FUSE HCD_EXACT PF_PAD Q8_PACKED Q6_PACKED MMVF_ROWS ATTN_LANECELL EXPERT_V2 TSUM LFUSE GDN_SPLIT QFUSE PLE_BATCH SH_STREAM ... strata serve: prompt 93792 tokens = 93757 reused + 35 read in 524 ms (66.8 tok/s), 57 generated in 1586 ms (36.0 tok/s), drafts accepted 42 of 61, 6 checkpoints strata serve: suffix drafts: 6 windows, 25 of 30 drafts accepted strata serve: prompt 93920 tokens = 93851 reused + 69 read in 890 ms (77.5 tok/s), 101 generated in 5713 ms (17.7 tok/s), drafts accepted 55 of 112, 6 checkpoints strata serve: suffix drafts: 8 windows, 26 of 40 drafts accepted strata serve: prompt 94648 tokens = 94022 reused + 626 read in 3100 ms (201.9 tok/s), 558 generated in 20890 ms (26.7 tok/s), drafts accepted 333 of 577, 6 checkpoints strata serve: suffix drafts: 26 windows, 65 of 128 drafts accepted strata serve: prompt 95244 tokens = 95205 reused + 39 read in 578 ms (67.5 tok/s), 52 generated in 1779 ms (29.2 tok/s), drafts accepted 37 of 46, 6 checkpoints strata serve: suffix drafts: 2 windows, 10 of 10 drafts accepted strata serve: no progress for 60 s during a request (reading the prompt (batched), done up to token 95573) - stopping the engine so the server starts it again (issue #29) strata serve: stall report (engine 0.1.40.3): stage "reading the prompt (batched), done up to token 95573" for 83 s; 0 layers served since the last finished step (0 = stopped, more = slow) expert pool: epoch 0, batch epoch 0: 0 of 0 jobs claimed, 0 done; 15 of 15 workers parked, 15 sleeping; mode 0 expert pool threads: w0=sleeping w1=sleeping w2=sleeping w3=sleeping w4=sleeping w5=sleeping w6=sleeping w7=sleeping w8=sleeping w9=sleeping w10=sleeping w11=sleeping w12=sleeping w13=sleeping w14=sleeping; host idle for 4906058 ms verify window (last window, not the current stage): 6 tokens at position 95294, host at layer step 1; the GPU rang 0; flags: served 1, plan (A) 0, copies (B) 0 threads waiting on the disk (state D): 0 of 41 memory: 49441 MiB resident, 401 MiB in swap, 6537 MiB RAM available; 801703 major page faults so far 2 s later: expert pool: epoch 0, batch epoch 0: 0 of 0 jobs claimed, 0 done; 15 of 15 workers parked, 15 sleeping; mode 0 expert pool threads: w0=sleeping w1=sleeping w2=sleeping w3=sleeping w4=sleeping w5=sleeping w6=sleeping w7=sleeping w8=sleeping w9=sleeping w10=sleeping w11=sleeping w12=sleeping w13=sleeping w14=sleeping; host idle for 4933694 ms verify window (last window, not the current stage): 6 tokens at position 95294, host at layer step 1; the GPU rang 0; flags: served 1, plan (A) 0, copies (B) 0 threads waiting on the disk (state D): 0 of 41 memory: 49419 MiB resident, 423 MiB in swap, 6559 MiB RAM available; 801708 major page faults so far strata: released the verify window's GPU waits (#267): the GPU finished in 0 ms ``` After restart session continued and even completed OK. Also suspicious: prefill speed = 300tps, ~2x slower that stock LLama.cpp at smaller contexts. In my setup 64 VRAM + 64GB GTT.
関連リンク
インストール・モデル・リリースへの站内リンク。