Issues / #1500

#1500 Engine stopped unexpectedly on strix halo

open · @stiff · 0 comments · View on GitHub

BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPModels & quantsSecurityWindows

Description

Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S 128k context, after doing ~50 turns session:

```
[strata] reading the prompt: 95,573 of 95,578 tokens, 112 s so far
[strata] reading the prompt: 95,573 of 95,578 tokens, 122 s so far
[strata] reading the prompt: 95,573 of 95,578 tokens, 132 s so far
[strata] reading the prompt: 95,573 of 95,578 tokens, 142 s so far
[strata] the engine stopped unexpectedly (exit code -6). The engine stopped itself because it had stopped making progress - a hang it caught. Its log line: strata serve: no progress for 60 s during a request (reading the prompt (batched), done up to token 95573) - stopping the engine so the server starts it again (issue #29) - please report it at github.com/Niko1221/Strata/issues. The next request starts the engine again. Its log: /home/stiff/tmp/strata/strata-iq3_s.log
```

ROCm 7.14 from AMD Debian repo, log snippet:

```
strata: gfx1151 (Strix Halo): 18 exact speed switches on by default (STRATA_GFX1151_DEFAULTS=0 turns them off; a switch you set is kept): GDN_HEAD GDN_PP GDN_CONVL2 GDN_NOY CVEC_FUSE HCD_EXACT PF_PAD Q8_PACKED Q6_PACKED MMVF_ROWS ATTN_LANECELL EXPERT_V2 TSUM LFUSE GDN_SPLIT QFUSE PLE_BATCH SH_STREAM
...
strata serve: prompt 93792 tokens = 93757 reused + 35 read in 524 ms (66.8 tok/s), 57 generated in 1586 ms (36.0 tok/s), drafts accepted 42 of 61, 6 checkpoints
strata serve: suffix drafts: 6 windows, 25 of 30 drafts accepted
strata serve: prompt 93920 tokens = 93851 reused + 69 read in 890 ms (77.5 tok/s), 101 generated in 5713 ms (17.7 tok/s), drafts accepted 55 of 112, 6 checkpoints
strata serve: suffix drafts: 8 windows, 26 of 40 drafts accepted
strata serve: prompt 94648 tokens = 94022 reused + 626 read in 3100 ms (201.9 tok/s), 558 generated in 20890 ms (26.7 tok/s), drafts accepted 333 of 577, 6 checkpoints
strata serve: suffix drafts: 26 windows, 65 of 128 drafts accepted
strata serve: prompt 95244 tokens = 95205 reused + 39 read in 578 ms (67.5 tok/s), 52 generated in 1779 ms (29.2 tok/s), drafts accepted 37 of 46, 6 checkpoints
strata serve: suffix drafts: 2 windows, 10 of 10 drafts accepted
strata serve: no progress for 60 s during a request (reading the prompt (batched), done up to token 95573) - stopping the engine so the server starts it again (issue #29)
strata serve: stall report (engine 0.1.40.3): stage "reading the prompt (batched), done up to token 95573" for 83 s; 0 layers served since the last finished step (0 = stopped, more = slow)
  expert pool: epoch 0, batch epoch 0: 0 of 0 jobs claimed, 0 done; 15 of 15 workers parked, 15 sleeping; mode 0
  expert pool threads: w0=sleeping w1=sleeping w2=sleeping w3=sleeping w4=sleeping w5=sleeping w6=sleeping w7=sleeping w8=sleeping w9=sleeping w10=sleeping w11=sleeping w12=sleeping w13=sleeping w14=sleeping; host idle for 4906058 ms
  verify window (last window, not the current stage): 6 tokens at position 95294, host at layer step 1; the GPU rang 0; flags: served 1, plan (A) 0, copies (B) 0
  threads waiting on the disk (state D): 0 of 41
  memory: 49441 MiB resident, 401 MiB in swap, 6537 MiB RAM available; 801703 major page faults so far
  2 s later:
  expert pool: epoch 0, batch epoch 0: 0 of 0 jobs claimed, 0 done; 15 of 15 workers parked, 15 sleeping; mode 0
  expert pool threads: w0=sleeping w1=sleeping w2=sleeping w3=sleeping w4=sleeping w5=sleeping w6=sleeping w7=sleeping w8=sleeping w9=sleeping w10=sleeping w11=sleeping w12=sleeping w13=sleeping w14=sleeping; host idle for 4933694 ms
  verify window (last window, not the current stage): 6 tokens at position 95294, host at layer step 1; the GPU rang 0; flags: served 1, plan (A) 0, copies (B) 0
  threads waiting on the disk (state D): 0 of 41
  memory: 49419 MiB resident, 423 MiB in swap, 6559 MiB RAM available; 801708 major page faults so far
strata: released the verify window's GPU waits (#267): the GPU finished in 0 ms
```

After restart session continued and even completed OK.

Also suspicious: prefill speed = 300tps, ~2x slower that stock LLama.cpp at smaller contexts. In my setup 64 VRAM + 64GB GTT.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.