Issues / #605

#605 --ple-io direct on rotational storage deadlocks prefill; the watchdog reports it as a generic engine stall (triage discriminator + --ple-io ram fix)

closed · @cjmckenna · 2 comentarios · En GitHub

BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

Descripción

### Summary

On a host whose model files sit on a 7200 RPM SATA disk, the default `--ple-io direct` deadlocks prompt reading. The 60 s watchdog kills the engine with a message that asks the user to file a bug, so the storage cause is invisible. `--ple-io ram` fixes it completely.

The point of this report is less "HDD is slow" and more that **this watchdog signature has many causes** (cf. #579 AMD/HIP, #251 verify-window, #224 PLE illegal access, #481 `SleepConditionVariableSRW`, #31 mid-generation freeze) and there is currently no way for a reporter to tell them apart. Below is the discriminator that worked, and a suggestion to put it in the stall report.

### Hardware

| | |
|---|---|
| GPU | RTX 3090 24 GB, sm_86, driver 595.91.07 |
| CPU | i9-10900F (AVX-2, no AVX-512) |
| RAM | 121 GB |
| Model storage | `/dev/sdb1`, **TOSHIBA DT01ACA2, `rotational=1`**, mq-deadline, `nr_requests=64` |
| Engine | 0.1.38, built locally, CUDA 13.3 |
| Model | Qwen3.8-Flash-Next IQ3_XXS |

Flags: `--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp rt --max-context 131072 --kv int8 --kv-resident 32768`

### Symptom

```
strata serve: no progress for 60 s during a request (reading the prompt (batched):
layer 1 of the prompt chunk from token 0) - stopping the engine so the server starts
it again (issue #29) - please report it at github.com/Niko1221/Strata/issues
```

```
stall report (engine 0.1.38): stage "reading the prompt (batched): layer 1 of the
prompt chunk from token 0" for 60 s; 0 layers served since the last finished step
  expert pool: epoch 0, batch epoch 0: 0 of 0 jobs claimed, 0 done; 9 of 9 workers parked
  host idle for 63272 ms
  verify window: 0 tokens at position 0, host at layer step 1; the GPU rang 0
```

Reproduced on every deliberate attempt with a fixed payload, including on freshly loaded engines immediately after a restart. I did not run a formally counted trial series, so I will not put a precise hit rate on it; it never once succeeded with that payload before the fix.

### What it is not

Each of these was ruled out by test, and most cost a rebuild or a full engine reload:

- **Not prompt size.** Stalls at 6,775 tokens; *succeeds* at 104K. Size is not the variable.
- **Not tool-schema count.** 53, 18 and 0 tools all stall on the same payload, at 81.6 s, 82.6 s and 81.5 s respectively.
- **Not concurrency.** A single sequential request reproduces it.
- **Not thinking or output budget.** `reasoning_effort=none` changes nothing.
- **Not the CUDA header/runtime mismatch of #542.** That was also present here and is genuinely worth fixing (rebuilt with `-DCUDAToolkit_ROOT=/usr/local/cuda-13.3`; note that `setup.py` passes only `-DCMAKE_CUDA_COMPILER`, so CMake independently resolved Debian's 12.4 `libcudart`). The stall survived the corrected build unchanged.
- **Not #577's unbuffered file tier.** `STRATA_UNBUFFERED_LOAD=0` measured about 1% *worse* here: median 43.4 s against a 42.9 s baseline, three cache-busted runs each on a fixed 27K-token prompt. Offered as a counter-data point since #577 is open; it presumably helps when the file cache is the hot path, which `--ple-io ram` removes.

### What it is: the discriminator

`/proc/<engine pid>/task/*/wchan` during the hang, identical across two dumps taken 10 s apart:

```
 16 wchan=blk_io_schedule      <-- state D, uninterruptible, blocked on block I/O
 16 wchan=futex_do_wait
  3 wchan=poll_schedule_timeout
```

Sixteen threads in `D` state on `blk_io_schedule` is the whole diagnosis. A CUDA-side deadlock parks threads on a futex or a CUDA sync, never on block I/O. It also explains `0 of 0 jobs claimed`: prefill never dispatches expert work because it is still waiting on PLE table rows.

`--ple-io direct` is documented as "unbuffered SSD reads". With `--ple-inflight` at its default of 256, against a 320M-row table, on a queue with `nr_requests=64`, a rotational actuator cannot service the random reads inside the watchdog window.

### Why it looks like a size limit, and is not

The most misleading part, and what cost the most time here:

| prompt | result |
|---|---|
| 42,548 chars of `"word word word..."` | **passes** |
| 21,274 chars of real prose and code | **stalls** |

Repetitive text hits a handful of n-gram rows that the row cache serves; token-diverse text scatters reads across the table. **Content diversity is the trigger, not length**, so a reporter who bisects on prompt size will never converge.

### Fix

`--ple-io ram`, locking the 27 GB table into RAM of 121 GB total:

| | before | after |
|---|---|---|
| 27,056-token prefill | never completes (watchdog kill) | **12 s, ~2,250 tok/s** |
| 103,374-token prompt | n/a | completes; decode 91 tok/s over 6,577 generated |
| stalls | every attempt | none in any run after the change |

### Suggestions

1. **Add I/O wait state to the stall report.** It already prints expert-pool and verify-window internals. A count of threads in `D` state plus their `wchan` would separate storage stalls from GPU and sync stalls at a glance, on every future report carrying this signature.
2. **Warn when the model's device is rotational.** `/sys/block/<dev>/queue/rotational` is a single read, and the engine already sizes the table and knows available RAM, so it could suggest or auto-select `--ple-io ram` when the table fits.
3. **Docs.** The description of `--ple-io direct` says "unbuffered SSD reads"; since it is the default, it seems worth stating outright that it is unsuitable for rotational storage.

Happy to run controlled follow-ups on this box, including `--ple-io mmap`, varied `--ple-inflight`, or an A/B with the model relocated to the SATA SSD.

En el sitio

Enlaces a install, modelos, releases.