Issues / #541
#541 AMD/HIP (R9700, engine 0.1.36): batched prompt read stalls at a repeatable token; STRATA_PLE_BATCH=0 avoids it
closed · @icodebot · 2 コメント · GitHub で見る
BenchmarksServer & APIAMD / HIPModels & quantsLinux
本文
## Summary On AMD (HIP, Radeon AI PRO R9700, gfx1201, Linux) engine 0.1.36 stalls while reading long prompts in the batched path; the 60 s watchdog (#29) catches it and restarts the engine. The same 223K-token prompt stalled twice at exactly the same position. With `STRATA_PLE_BATCH=0` (the workaround from #224) the same prompt was read end to end. Possibly the same family as #217 / #251 / #224, but on AMD/HIP and on a newer engine than those fixes. ## Environment - Strata `36fa455`, engine 0.1.36, backend `hip` - GPU: AMD Radeon AI PRO R9700 32 GB (gfx1201), ROCm 7.2.4, power cap 210 W (raised to 300 W partway through the successful run below) - CPU: Ryzen 9 9950X3D, 64 GB RAM (60 GB usable) - OS: CachyOS (Arch), kernel 7.2.8-1-cachyos - Model: Coder IQ1_M (`coder-iq1_m`), `--prefill auto` (8192-token chunks), `--max-context 262144`, `--kv int8`, `--kv-resident 32768`, `--expert-cache auto` (all 12288 experts resident in VRAM), MTP draft head on - Client: OpenCode (OpenAI-compatible API) ## What happened Five stalls in one log, all in the batched prompt read: ``` strata serve: no progress for 60 s during a request (reading the prompt (batched): finishing the chunk from token 16633) - stopping the engine so the server starts it again (issue #29) strata serve: no progress for 60 s during a request (reading the prompt (batched): finishing the chunk from token 9040) - stopping the engine so the server starts it again (issue #29) strata serve: no progress for 60 s during a request (reading the prompt (batched): finishing the chunk from token 213840) - stopping the engine so the server starts it again (issue #29) strata serve: no progress for 60 s during a request (reading the prompt (batched): finishing the chunk from token 205648) - stopping the engine so the server starts it again (issue #29) strata serve: no progress for 60 s during a request (reading the prompt (batched): finishing the chunk from token 205648) - stopping the engine so the server starts it again (issue #29) ``` - The last two were the same request (an OpenCode "compact" of a 223,345-token session), retried; both stalled at token 205648. - Stalls also happened at small offsets (9040, 16633), so it is not only a long-context issue. - Because the restart drops the KV cache, every retry re-reads the whole prompt, so a long session could not be compacted at all. ## Workaround that worked Restarted the server with `STRATA_PLE_BATCH=0` (checked it reached the engine via `/proc/<pid>/environ`). The same compaction prompt then completed: ``` strata serve: prompt 223345 tokens = 0 reused + 223345 read in 358145 ms (623.6 tok/s), 2860 generated in 48035 ms (59.5 tok/s), drafts accepted 1729 of 2610, 6 checkpoints ``` Prompt speed with the workaround (~600-635 tok/s) was about the same as before it (~630 tok/s on the earlier attempts). One success after two failures at the same spot; I have not run it more times. The full log (`strata-coder-iq1_m.log`, including the stall reports) is available if useful.
関連リンク
インストール・モデル・リリースへの站内リンク。