Issues / #541

#541 AMD/HIP (R9700, engine 0.1.36): batched prompt read stalls at a repeatable token; STRATA_PLE_BATCH=0 avoids it

closed · @icodebot · 2 commentaires · Sur GitHub

BenchmarksServer & APIAMD / HIPModels & quantsLinux

Description

## Summary

On AMD (HIP, Radeon AI PRO R9700, gfx1201, Linux) engine 0.1.36 stalls while reading long prompts in the batched path; the 60 s watchdog (#29) catches it and restarts the engine. The same 223K-token prompt stalled twice at exactly the same position. With `STRATA_PLE_BATCH=0` (the workaround from #224) the same prompt was read end to end.

Possibly the same family as #217 / #251 / #224, but on AMD/HIP and on a newer engine than those fixes.

## Environment

- Strata `36fa455`, engine 0.1.36, backend `hip`
- GPU: AMD Radeon AI PRO R9700 32 GB (gfx1201), ROCm 7.2.4, power cap 210 W (raised to 300 W partway through the successful run below)
- CPU: Ryzen 9 9950X3D, 64 GB RAM (60 GB usable)
- OS: CachyOS (Arch), kernel 7.2.8-1-cachyos
- Model: Coder IQ1_M (`coder-iq1_m`), `--prefill auto` (8192-token chunks), `--max-context 262144`, `--kv int8`, `--kv-resident 32768`, `--expert-cache auto` (all 12288 experts resident in VRAM), MTP draft head on
- Client: OpenCode (OpenAI-compatible API)

## What happened

Five stalls in one log, all in the batched prompt read:

```
strata serve: no progress for 60 s during a request (reading the prompt (batched): finishing the chunk from token 16633) - stopping the engine so the server starts it again (issue #29)
strata serve: no progress for 60 s during a request (reading the prompt (batched): finishing the chunk from token 9040) - stopping the engine so the server starts it again (issue #29)
strata serve: no progress for 60 s during a request (reading the prompt (batched): finishing the chunk from token 213840) - stopping the engine so the server starts it again (issue #29)
strata serve: no progress for 60 s during a request (reading the prompt (batched): finishing the chunk from token 205648) - stopping the engine so the server starts it again (issue #29)
strata serve: no progress for 60 s during a request (reading the prompt (batched): finishing the chunk from token 205648) - stopping the engine so the server starts it again (issue #29)
```

- The last two were the same request (an OpenCode "compact" of a 223,345-token session), retried; both stalled at token 205648.
- Stalls also happened at small offsets (9040, 16633), so it is not only a long-context issue.
- Because the restart drops the KV cache, every retry re-reads the whole prompt, so a long session could not be compacted at all.

## Workaround that worked

Restarted the server with `STRATA_PLE_BATCH=0` (checked it reached the engine via `/proc/<pid>/environ`). The same compaction prompt then completed:

```
strata serve: prompt 223345 tokens = 0 reused + 223345 read in 358145 ms (623.6 tok/s), 2860 generated in 48035 ms (59.5 tok/s), drafts accepted 1729 of 2610, 6 checkpoints
```

Prompt speed with the workaround (~600-635 tok/s) was about the same as before it (~630 tok/s on the earlier attempts). One success after two failures at the same spot; I have not run it more times.

The full log (`strata-coder-iq1_m.log`, including the stall reports) is available if useful.

Sur le site

Liens install, modèles, releases.