Issues / #605
#605 --ple-io direct on rotational storage deadlocks prefill; the watchdog reports it as a generic engine stall (triage discriminator + --ple-io ram fix)
closed · @cjmckenna · 2 comments · View on GitHub
BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
Description
### Summary On a host whose model files sit on a 7200 RPM SATA disk, the default `--ple-io direct` deadlocks prompt reading. The 60 s watchdog kills the engine with a message that asks the user to file a bug, so the storage cause is invisible. `--ple-io ram` fixes it completely. The point of this report is less "HDD is slow" and more that **this watchdog signature has many causes** (cf. #579 AMD/HIP, #251 verify-window, #224 PLE illegal access, #481 `SleepConditionVariableSRW`, #31 mid-generation freeze) and there is currently no way for a reporter to tell them apart. Below is the discriminator that worked, and a suggestion to put it in the stall report. ### Hardware | | | |---|---| | GPU | RTX 3090 24 GB, sm_86, driver 595.91.07 | | CPU | i9-10900F (AVX-2, no AVX-512) | | RAM | 121 GB | | Model storage | `/dev/sdb1`, **TOSHIBA DT01ACA2, `rotational=1`**, mq-deadline, `nr_requests=64` | | Engine | 0.1.38, built locally, CUDA 13.3 | | Model | Qwen3.8-Flash-Next IQ3_XXS | Flags: `--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp rt --max-context 131072 --kv int8 --kv-resident 32768` ### Symptom ``` strata serve: no progress for 60 s during a request (reading the prompt (batched): layer 1 of the prompt chunk from token 0) - stopping the engine so the server starts it again (issue #29) - please report it at github.com/Niko1221/Strata/issues ``` ``` stall report (engine 0.1.38): stage "reading the prompt (batched): layer 1 of the prompt chunk from token 0" for 60 s; 0 layers served since the last finished step expert pool: epoch 0, batch epoch 0: 0 of 0 jobs claimed, 0 done; 9 of 9 workers parked host idle for 63272 ms verify window: 0 tokens at position 0, host at layer step 1; the GPU rang 0 ``` Reproduced on every deliberate attempt with a fixed payload, including on freshly loaded engines immediately after a restart. I did not run a formally counted trial series, so I will not put a precise hit rate on it; it never once succeeded with that payload before the fix. ### What it is not Each of these was ruled out by test, and most cost a rebuild or a full engine reload: - **Not prompt size.** Stalls at 6,775 tokens; *succeeds* at 104K. Size is not the variable. - **Not tool-schema count.** 53, 18 and 0 tools all stall on the same payload, at 81.6 s, 82.6 s and 81.5 s respectively. - **Not concurrency.** A single sequential request reproduces it. - **Not thinking or output budget.** `reasoning_effort=none` changes nothing. - **Not the CUDA header/runtime mismatch of #542.** That was also present here and is genuinely worth fixing (rebuilt with `-DCUDAToolkit_ROOT=/usr/local/cuda-13.3`; note that `setup.py` passes only `-DCMAKE_CUDA_COMPILER`, so CMake independently resolved Debian's 12.4 `libcudart`). The stall survived the corrected build unchanged. - **Not #577's unbuffered file tier.** `STRATA_UNBUFFERED_LOAD=0` measured about 1% *worse* here: median 43.4 s against a 42.9 s baseline, three cache-busted runs each on a fixed 27K-token prompt. Offered as a counter-data point since #577 is open; it presumably helps when the file cache is the hot path, which `--ple-io ram` removes. ### What it is: the discriminator `/proc/<engine pid>/task/*/wchan` during the hang, identical across two dumps taken 10 s apart: ``` 16 wchan=blk_io_schedule <-- state D, uninterruptible, blocked on block I/O 16 wchan=futex_do_wait 3 wchan=poll_schedule_timeout ``` Sixteen threads in `D` state on `blk_io_schedule` is the whole diagnosis. A CUDA-side deadlock parks threads on a futex or a CUDA sync, never on block I/O. It also explains `0 of 0 jobs claimed`: prefill never dispatches expert work because it is still waiting on PLE table rows. `--ple-io direct` is documented as "unbuffered SSD reads". With `--ple-inflight` at its default of 256, against a 320M-row table, on a queue with `nr_requests=64`, a rotational actuator cannot service the random reads inside the watchdog window. ### Why it looks like a size limit, and is not The most misleading part, and what cost the most time here: | prompt | result | |---|---| | 42,548 chars of `"word word word..."` | **passes** | | 21,274 chars of real prose and code | **stalls** | Repetitive text hits a handful of n-gram rows that the row cache serves; token-diverse text scatters reads across the table. **Content diversity is the trigger, not length**, so a reporter who bisects on prompt size will never converge. ### Fix `--ple-io ram`, locking the 27 GB table into RAM of 121 GB total: | | before | after | |---|---|---| | 27,056-token prefill | never completes (watchdog kill) | **12 s, ~2,250 tok/s** | | 103,374-token prompt | n/a | completes; decode 91 tok/s over 6,577 generated | | stalls | every attempt | none in any run after the change | ### Suggestions 1. **Add I/O wait state to the stall report.** It already prints expert-pool and verify-window internals. A count of threads in `D` state plus their `wchan` would separate storage stalls from GPU and sync stalls at a glance, on every future report carrying this signature. 2. **Warn when the model's device is rotational.** `/sys/block/<dev>/queue/rotational` is a single read, and the engine already sizes the table and knows available RAM, so it could suggest or auto-select `--ple-io ram` when the table fits. 3. **Docs.** The description of `--ple-io direct` says "unbuffered SSD reads"; since it is the default, it seems worth stating outright that it is unsuitable for rotational storage. Happy to run controlled follow-ups on this box, including `--ple-io mmap`, varied `--ple-inflight`, or an A/B with the model relocated to the SATA SSD.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.