Issues / #1425
#1425 [Windows, RTX 5090 32 GB, NVFP4 fork] ple: --ple-io direct streams the n-gram table at ~34 MB/s while the SSD does 1.4 GB/s - fresh long prompts prefill ~5x slow; cold passes can trip the #29 watchdog
open · @RichardFreml · 0 comentários · No GitHub
BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Descrição
## Summary
On Windows, layer 1's **PLE (n-gram) table** is read from the SSD on every fresh prompt. With the
default `--ple-io direct`, **unbuffered** reads only reach **~34 MB/s** while the same SSD does
**1.4–1.7 GB/s** sequentially. Because that I/O is on the prompt path, a real (diverse) conversation
prefills at **~1,240 tok/s** while a synthetic filler prompt of the same length prefills at
**~6,500–7,300 tok/s** (its n-grams repeat → cached rows → no I/O), with the GPU idle during the stalls.
Switching to **`--ple-io mmap`** fixes the warm case (**6,037 tok/s, disk 0**) because the OS page
cache holds the table pages across requests, but the **first cold pass still streams at ~156 MB/s**
and can trip the `#29` no-progress watchdog (engine self-terminates, request errors, server reloads).
## Environment
- **Engine:** `0.1.40-nvfp4.3`, fork [`sergqwer/strata-nvfp4`](https://github.com/sergqwer/strata-nvfp4)
(based on upstream **0.1.40**), backend cuda, arch `sm_120`.
- **GPU:** RTX 5090 32 GB (1×, `--layer-split` n/a).
- **Host:** AMD Ryzen Threadripper 3970X (32C/64T, Zen 2, **no AVX-512**), **128 GB RAM**, Windows.
- **Disk:** GIGABYTE GP-GSM2NE3100TNTD NVMe 1 TB — **~1,400–1,670 MB/s sequential** (observed at model load).
- **Model:** Qwen3.8-Flash-Next NVFP4 pack; `models/ple-fp8.gguf` = **47.68 GiB** (320,001,536 × 160 n-gram table).
- **Config:** `parallel: 2` × 524,288 ctx, YaRN ×2, `--kv int8 --kv-resident 32768`,
`--ple-gguf models/ple-fp8.gguf`, `--batch-mtp`, `--conversation-cache-mib 0`.
- **`--ple-io`:** default (`direct`) unless stated.
## Reproduction
1. Use (or restart) the engine and send a **fresh, real long prompt** whose n-grams have never been
read — e.g. a ~168k-token coding-agent session history. (A synthetic *repetitive* filler will
**not** reproduce; see measurements.)
2. Watch disk read (Windows: `psutil.disk_io_counters()`, e.g. `disk_read_mb` in `/metrics`).
3. With **`--ple-io direct`**: during the prompt pass CPU ~3 %, GPU **0–4 %**, power ~27 W, and
**disk read ~34 MB/s sustained**; prefill ~1,240 tok/s. Periodically the request hangs and the
engine aborts: `no progress for 60 s … (reading the prompt (batched): layer 1 of the prompt chunk
from token 8248)` → exit `3221226505`.
4. With **`--ple-io mmap`**: a **warm** repeat of the same session runs `disk=0`, GPU 100 %,
**6,037 tok/s**. A **cold** pass streams at ~156 MB/s and can still trip the watchdog.
## Expected Behavior
A fresh long prompt should prefill at roughly the same order as a cached/repetitive one, and the
no-progress watchdog should not fire on I/O-bound prompt reads.
## Actual Behavior
- `--ple-io direct`: PLE cold reads ~**34 MB/s** (unbuffered), GPU idle → prefill **~1,240 tok/s**
on a 168k real session vs **~7,300 tok/s** for filler.
- `--ple-io mmap`: warm **6,037 tok/s** (disk 0), but cold ~**156 MB/s** and a 60 s no-progress abort.
- Same disk does **1.4–1.7 GB/s** sequentially (model load), so **the disk is not the limit — the
read mode is**.
## Measurements (7. 10. 2026)
| scenario | prefill | disk read |
|---|---:|---:|
| real 168k session, `--ple-io direct` | 1,240 tok/s | 34 MB/s (3.8 GB over the pass) |
| real 168k session, `--ple-io mmap`, **warm cache** | **6,037 tok/s** | **0** |
| synthetic 233k fresh prompt, `--ple-io mmap`, **cold** | 1,725 tok/s | 156 MB/s (14.6 GB) |
| synthetic filler (repetitive n-grams) | 6,500–7,300 tok/s | 0 |
Per-second `psutil` sample during a `direct` stall (state=`reading`):
```
cpu=3.0% gpu_util=0% gpu_power=27W disk_read=34.7 MB/s <- held for ~10 s per chunk
...
gpu_util=100% gpu_power=316W pcie_rx=46 GB/s disk_read=0.03 <- then the compute burst
```
Stall report (engine 0.1.40), same class as #29:
```
strata serve: no progress for 60 s during a request (reading the prompt (batched):
layer 1 of the prompt chunk from token 8248) - stopping the engine ...
expert pool: 0 of 0 jobs claimed, 31 of 31 workers parked, 31 sleeping
host idle for 133717 ms
wrote the thread stacks to ...strata-stall-8592.dmp
```
## Analysis
`docs/NVFP4.md` says layer 1 adds **16 rows of the n-gram table per token** and the table *"stays on
the SSD (16 page reads a token either way)"*, and lists the PLE read path as **"plan v0.3 P2"**
(`--ple-io direct|mmap|ram`). On this Windows host `direct` (= unbuffered) is pathologically slow
(~2.4 % of the disk's sequential rate), while `mmap` reaches ~156 MB/s cold and ~0 MB/s warm.
Upstream already carries **"unbuffered loads on Windows (#357, #362)"**, so this may be a known
class. Note `--ple-io ram` (lock the whole table in RAM) is documented **Linux/macOS only** — a
Windows box with 128 GB can't use it, and the table (47.68 GiB) would not fit the ~32 GiB free anyway.
## Additional Context
- **Workaround:** `--ple-io mmap` — warm sessions ~6,000 tok/s, disk 0. Documented locally.
- **A/B `--ple-row-cache`** (8M rows = 720 MB vs default 1M): **no difference**
(cold/warm 2,985/6,788 vs 3,230/6,784) → the OS page cache does the work, row-cache is redundant.
- **Related issues:** #1407 (expert **file tier** streaming, same watchdog #29 class; Linux/AMD),
#1341 (`--pipeline-windows` >100k deadlock, #29), #29 (the watchdog), #357/#362 (unbuffered Windows).
## Questions
1. Is ~34 MB/s expected for `--ple-io direct` on Windows, or is the unbuffered path a bug there?
2. Should an I/O-bound prompt read be able to trip the 60 s no-progress watchdog (i.e. is I/O time
counted as "no progress")?
3. Is Windows support for `--ple-io ram` (or an equivalent "keep the hot table region in RAM") planned?
No site
Links install, modelos, releases.