Issues / #1341
#1341 serve: a >100k-token prompt deadlocks the pipelined verify-window reader (--pipeline-windows >= 1) - the #29 watchdog aborts the engine and the server reloads the model
open · @btc1000w · 2 コメント · GitHub で見る
BenchmarksServer & APIMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows
本文
## Summary
On engine 0.1.40 (server 0.1.40.1) with a dual-GPU layer split, a prompt of about 120k tokens hangs while "reading the prompt". The engine's own #29 watchdog fires after 60 s, writes a stall dump, and aborts the engine; the server then reloads the model (about 2 minutes) and the request fails. Setting `--pipeline-windows 0` makes the same long cold reads pass.
The hang is in the **pipelined verify-window prompt reader** (`read_windows_pl`), which is only used when `--pipeline-windows >= 1`. The engine prints no `illegal memory access` in this reproduction.
## Environment
- engine 0.1.40, server 0.1.40.1 (82f46a8), Windows 11, CUDA 13.0 release build, compute capability 8.6
- 2x RTX Laptop (3080 + 3070), 16384 MiB each, driver 616.56; CPU Intel Xeon E5-2680 v4
- model Qwen3.8-Flash-Next IQ3_XXS (24576 experts, 48 layers)
- args: `--pack ... --native ... --ple-gguf ... --expert-profile ... --expert-cache auto --prefill auto:32768 --spec 4 --spec-min-p 0.70 --mtp ... --max-context 262144 --kv int8 --kv-resident 32768 --vision --vram-reserve-mib 690 --pool-workers 8 --trim-stage-weights --pipeline-windows 2 --layer-split 26`
- 630 MiB of VRAM free with everything loaded; about 34 GB of RAM free -> not an OOM.
## Observed
Production, same config: a request with a 129,038-token prompt failed (finish=error, 0 output tokens) and the server reloaded the engine. Earlier the same day a 107,882-token prompt ended with `prefill copy_i32: an illegal memory access was encountered` (the sticky-label class) plus the #29 watchdog, in the **batched** reader (`stage "reading the prompt (batched): layer 27 of the prompt chunk from token 107882"`).
Reproduced 2/2 with a cold ~120,000-token prompt (a fresh, uncached prompt; nothing reused):
```
strata serve: no progress for 60 s during a request (reading the prompt (pipelined verify windows), from token 120013) - stopping the engine so the server starts it again (issue #29)
strata serve: stall report (engine 0.1.40): stage "reading the prompt (pipelined verify windows), from token 120013" for 57 s; 0 layers served since the last finished step (0 = stopped, more = slow)
expert pool: epoch 0, batch epoch 0: 0 of 0 jobs claimed, 0 done; 8 of 8 workers parked, 8 sleeping; mode 0
wrote the thread stacks to C:\strata\Strata\strata-stall-16932.dmp (attach it to the issue)
strata: released the verify window's GPU waits (#267): the GPU finished in 0 ms
strata: released the verify window's GPU waits (#267): the GPU finished in 0 ms
strata: released the verify window's GPU waits (#267): the GPU finished in 0 ms
strata: released the verify window's GPU waits (#267): the GPU finished in 0 ms
strata: released the verify window's GPU waits (#267): the GPU did not finish within 5 s
strata: released the verify window's GPU waits (#267): the GPU did not finish within 5 s
strata: released the verify window's GPU waits (#267): the GPU finished in 20 ms
strata: released the verify window's GPU waits (#267): the GPU finished in 20 ms
```
The second run was identical (`for 56 s`, dump `strata-stall-22140.dmp`). The prompt read rate was normal before the hang (120k tokens in about 95 s, ~1260 tok/s), so it is not a slow-read timeout.
Note the expert pool line: **0 of 0 jobs claimed** -> no prefill job is running at all; the reader is parked waiting on something that is never signalled, not merely slow.
Instrumentation: the two reproduction runs set `CUDA_LAUNCH_BLOCKING=1` and `STRATA_PF_STEP_SYNC=1` to localise the stage. The production failures ran with the same config and no instrumentation. Launch blocking normally makes races *less* likely, which points to a deterministic missing-signal deadlock rather than a timing race.
## Suspected cause
The stage label `reading the prompt (pipelined verify windows)` is emitted only from `read_windows_pl` (src/program/generate.cpp:8521; the label is at :8522). That reader is selected only when `--pipeline-windows >= 1`:
```cpp
// src/program/generate.cpp:8608
if (pipe && pl_pw >= 1 && std::getenv("STRATA_LOGPOS") == nullptr) return read_windows_pl(a, b, e);
```
with `int pl_pw = pipe ? o.pipeline_windows : 0;` (generate.cpp:8494). So the deadlock is in the overlapped cross-stage verify-window prompt read. The engine does not report an illegal access in this reproduction, so the `prefill copy_i32` label seen in production is the sticky `cudaGetLastError()` mis-attribution (see src/kernels/cuda/qsa.cu:77) rather than the actual faulting kernel.
The earlier production signature (batched reader, layer 27) may be a sibling manifestation of the same verify-window handshake; I could not attribute it separately.
## Workaround (verified)
`--pipeline-windows 0`: 3/3 cold long reads pass, no fault, no stall, no reload:
| prompt tokens | result | time | cached_tokens |
|---|---|---|---|
| 121,027 | HTTP 200 | 64.8 s | 0 |
| 126,027 | HTTP 200 | 68.7 s | 0 |
| 131,026 | HTTP 200 | 70.8 s | 0 |
(all genuinely cold; `--prefill auto:32768` unchanged). Before the change: 2/2 cold ~120k reads stalled.
Trade-off: this gives up the ~+16-21% decode and ~129 MiB of headroom that `--pipeline-windows` provides (docs/MULTI_GPU.md).
## Dumps
`strata-stall-16932.dmp`, `strata-stall-22140.dmp` (this reproduction), `strata-stall-23716.dmp` (production, batched path, layer 27 / token 107882). Happy to attach any of them.
## Related
- #29 (engine watchdog), #267 (verify-window GPU-wait release), #964 (open: dual-GPU + >100k prompt + illegal access + verify-window hang), #1275.
関連リンク
インストール・モデル・リリースへの站内リンク。