Issues / #1629
#1629 [Linux, RTX 3090, stock 0.1.40.2] --ple-io direct on a DRAM-less NVMe reads the PLE table at 3.7-8.3 tok/s on the prompt path and trips the #29 watchdog; --ple-io mmap is 40-100x faster, first cold pass included
open · @bluebirdlboro · 0 コメント · GitHub で見る
BenchmarksServer & APINVIDIA / CUDAModels & quantsWindowsLinux
本文
## Summary
On a **stock 0.1.40.2** engine (Linux/CUDA, RTX 3090 24 GB), the default `--ple-io direct` reads layer 1's
n-gram (PLE) table so slowly from a **DRAM-less NVMe SSD** that a real coding-agent prompt never gets past
layer 1: `0 layers served` in 60 s, so the #29 no-progress watchdog stops the engine (client sees
`the engine stopped unexpectedly (exit code -6)`), and the server reloads it (~3 min cold start).
The prompt path is the whole story: with `--ple-io direct`, a **cold 59-token prompt takes 16.1 s to read**
(3.7 tok/s) and a 170-token prompt 20.4 s (8.3 tok/s) — no prefill work involved, just the PLE row reads.
Switching to **`--ple-io mmap`** makes the same reads 293–508 tok/s (40–100×) and, unlike the first cold
pass in #1425, it holds on the **first** pass: a 38,808-token prompt read in 101 s with steady progress,
no watchdog trip.
This is the Linux/CUDA/stock-engine counterpart to #1425 (Windows, NVFP4 fork: `direct` ≈ 34 MB/s), and the
case where `--ple-io ram` (#746, #605) is **not available** because the host cannot spare 28.8 GB.
## Environment
- **Engine:** stock `0.1.40.2` (`strata serve: stall report (engine 0.1.40.2)`), served via the pack's own
`serve/server.py --engine strata`; backend cuda.
- **GPU:** RTX 3090 24 GB (1×), `--vram-reserve-mib 700`.
- **Host:** Ubuntu, 62 GiB RAM, 9 GiB swap (≈5 GiB already used); ~47 GiB resident for the engine.
- **Disk:** WD Blue SN580 1 TB NVMe — `/sys/block/nvme0n1/queue/rotational` = `0`, **DRAM-less** (no host
DRAM for the FTL map).
- **Model:** `Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS` (two shards) + `--ple-gguf` shard **28.8 GB**,
`--expert-profile … --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp … --max-context 131072
--kv int8 --vision`. Same result observed with the IQ3_S pack.
- **`--ple-io`:** default (`direct`) unless stated; `--ple-row-cache` / `--ple-inflight` left at defaults
(1 048 576 / 256) in both arms.
## Reproduction
Fresh (diverse) prompt of agent-session size through the OpenAI-compatible endpoint, on a freshly started
engine so nothing of that prompt is cached:
- `--ple-io direct`: engine log goes to
```
strata serve: no progress for 60 s during a request (reading the prompt (batched): layer 1 of the prompt chunk from token 0) - stopping the engine so the server starts it again (issue #29)
strata serve: stall report (engine 0.1.40.2): stage "reading the prompt (batched): layer 1 of the prompt chunk from token 0" for 60 s; 0 layers served since the last finished step
threads waiting on the disk (state D): 2 of 13 (rq_qos_wait x2) - the engine is waiting for the drive: a rotational or failing disk with --ple-io direct (#605: --ple-io ram), or a drive too slow for the reads asked of it
memory: 47676 MiB resident, 1485 MiB in swap, 13493 MiB RAM available; 305019 major page faults so far
```
Reproduced on every attempt with an 18 672-token prompt and a 30 731-token prompt; small prompts
(< 200 tokens) finish but are also slow to read.
- `--ple-io mmap`: same prompts, same engine, one flag changed — all pass (numbers below).
## Measurements
Cold prompt reads (no prefix reuse), from the engine's own `prompt … read in … ms` lines:
| prompt | `--ple-io direct` | `--ple-io mmap` |
|---|---|---|
| 59 tokens | 16 145 ms (3.7 tok/s) | — |
| 96–170 tokens | 11 838 / 20 374 ms (7.3 / 8.3 tok/s) | — |
| 18 672 tokens | never posted (deadline hit at layer 1) | **36 780 ms (507.7 tok/s)** |
| 30 731 tokens | never posted | **104 818 ms (293.2 tok/s)** |
| 38 808 tokens | not attempted | **101 255 ms (383.3 tok/s)** |
End-to-end through the client path (same prompt size, both arms):
- `direct`: request fails — `the engine stopped unexpectedly (exit code -6)`; server restarts the engine
(`experts loaded: 39.97 GiB at 0.83 GiB/s (98 s)`), ~3 min before it can serve again.
- `mmap`: 18 672-token agent prompt answers in 73 s wall clock (prompt read 36.8 s + 486 tokens decoded at
48.7 tok/s); 30 731-token prompt completes with exit code 0. Prefix reuse stays cheap afterwards
(`38803 reused + 5 read in 84 ms`).
## Why `--ple-io ram` is not an option here
`--ple-io ram` needs the PLE shard resident (**28.8 GB**), but the host has 62 GiB total with ~47 GiB already
resident for the engine (experts in RAM) and ~13 GiB free — it would push the machine into swap (swap was
already 5 GiB used during the failing run). `mmap` works because the touched row set of a prompt is far
smaller than the whole table and the page cache keeps it across requests, so it costs little RAM.
## The stall-report hint is misleading for this class of disk
`a rotational or failing disk with --ple-io direct` sent me looking at the drive first. A DRAM-less NVMe is
neither rotational nor failing — it does 1.4 GB/s sequentially — but random 4 K reads of ~90-byte rows pay an
extra FTL metadata read each, which is exactly what `direct` (O_DIRECT, no OS cache) does on the prompt path.
Suggest mentioning "DRAM-less SSD / random small reads" (and `--ple-io mmap`) in that hint.
## What would help
1. Should `--ple-io mmap` be the documented recommendation for DRAM-less NVMe when `ram` is unaffordable?
(It is the only setting that works here, and it holds on the cold pass — unlike the 156 MB/s cold pass
reported in #1425.)
2. The 60 s watchdog has no slack for a *slow but progressing* read: on `direct` it fires at layer 1 with
nothing served. Is making it configurable, or exempting the prompt-read stage, on the table?
3. Would raising `--ple-row-cache` / `--ple-inflight` (defaults 1 048 576 / 256) help in this regime? I have
not tried them; the batched/read-ahead PLE read in the style of #1323 sounds like the real fix.
4. Is `--ple-io ram` meant to be the answer whenever RAM allows (per #746/#605) and `mmap` the fallback, or is
there a third mode planned (e.g. buffered reads with read-ahead) for hosts that can afford neither the
RAM nor the O_DIRECT penalty?
Related: #1425 (Windows/NVFP4 fork, `direct` ≈ 34 MB/s), #1211 (`ram` null result on fast NVMe),
#605 / #746 (`ram` as the storage lever), #1323 (batched unbuffered reads), #29 (the watchdog).
関連リンク
インストール・モデル・リリースへの站内リンク。