Pull requests / #382

#382 HIP: MTP prompt pass per group by default (fixes prompt hang on gfx1201 with KV streaming)

closed · @abhinand5 · 0 コメント · GitHub で見る

BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsLinux

本文

I've been running the full GSQ-RCO IQ3_XXS on a Radeon AI PRO R9700 (gfx1201, ROCm 7.2.4, Linux) with a 131K context, and long prompts would sometimes just stop partway through: the stall report said `reading the prompt (batched) ... 0 layers served`, the GPU sat at 100%, and after the watchdog fired the server restarted the engine and the request failed. It happened roughly once per round of a mixed load (chats plus a fresh 32K and a fresh 7K prompt), so it was very noticeable with an agent sending long contexts.

**What's going on**

Since `ptrace_scope=1` doesn't allow attaching, I ran the engine as a child of gdb and let the watchdog's `abort()` dump every thread. Every capture (four on 0.1.25, one on 0.1.31) had the main thread inside `MtpDrafter::prefill`, blocked in `hipGraphLaunch` or in the final `cudaStreamSynchronize`, while the stager, the expert pool and the verify flags were all idle.

On 0.1.31 I expected E-9 to take over, but `Prefill::draft_kv` declines whenever the drafter's KV isn't `kv_mode == 0`. With KV streaming on (`--kv-resident 32768`, which is what lets 131K fit next to a 25 GiB expert cache), the drafter's K/V is a ring (`kv_mode == 2`), so every chunk falls back to the E-4 path. That path queues a graph launch plus four device copies per group, around 2,000 groups per 8192-token chunk, with one sync at the end, and on this card that queue sometimes never completes. I couldn't pin down the exact HIP-level reason (no umr/rocgdb here), but the older per-group loop (`STRATA_MTP_PREFILL_SYNC=1`) never hangs.

**The change**

HIP builds now default to the per-group loop. `STRATA_MTP_PREFILL_SYNC=0` brings E-4 back, `=1` still forces the old loop, and CUDA builds behave exactly as before. I also made the variable read its value: until now any value, including `0`, turned the old loop on.

**Testing** (R9700 32 GB, Ryzen 7 9800X3D, 30 GB RAM, full IQ3_XXS native pack with `--mmap-experts`, `--kv int8 --kv-resident 32768 --max-context 131072 --spec 4 --mtp`)

Each round is a short code chat, a short prose chat, a fresh 32K prompt, a code turn at 32K depth, a fresh 7K prompt and a code turn at 7K depth, all through `serve/server.py`.

| build | rounds | hangs |
|---|---|---|
| 0.1.31 as released | 1 and 1 | hung in round 1 both times |
| 0.1.31 + `STRATA_MTP_PREFILL_SYNC=1` | 30 | 0 |
| this branch, no env var | 30 | 0 |

The per-group loop doesn't cost anything I can measure. On this branch the median warm 32K prompt reads at 1,131 tok/s and a 7K one at 1,153 tok/s (decode 72-80 tok/s on code). On 0.1.25, where E-4 survived long enough to compare, both paths read 32K prompts at about 1,050-1,100 tok/s. `ctest` passes 35/35 with `hip_prefill_hipblaslt_gemm` skipped, as usual.

A longer-term fix might be letting E-9 write the drafter's ring as well, so this fallback isn't hit at all with KV streaming; I left that alone since it's your design.

---
*This PR was drafted with AI assistance (Claude), and tested on my machine.*

🤖 Generated with [Claude Code](https://claude.com/claude-code)

関連リンク

インストール・モデル・リリースへの站内リンク。