Issues / #1629

#1629 [Linux, RTX 3090, stock 0.1.40.2] --ple-io direct on a DRAM-less NVMe reads the PLE table at 3.7-8.3 tok/s on the prompt path and trips the #29 watchdog; --ple-io mmap is 40-100x faster, first cold pass included

open · @bluebirdlboro · 0 comments · View on GitHub

BenchmarksServer & APINVIDIA / CUDAModels & quantsWindowsLinux

Description

## Summary

On a **stock 0.1.40.2** engine (Linux/CUDA, RTX 3090 24 GB), the default `--ple-io direct` reads layer 1's
n-gram (PLE) table so slowly from a **DRAM-less NVMe SSD** that a real coding-agent prompt never gets past
layer 1: `0 layers served` in 60 s, so the #29 no-progress watchdog stops the engine (client sees
`the engine stopped unexpectedly (exit code -6)`), and the server reloads it (~3 min cold start).

The prompt path is the whole story: with `--ple-io direct`, a **cold 59-token prompt takes 16.1 s to read**
(3.7 tok/s) and a 170-token prompt 20.4 s (8.3 tok/s) — no prefill work involved, just the PLE row reads.
Switching to **`--ple-io mmap`** makes the same reads 293–508 tok/s (40–100×) and, unlike the first cold
pass in #1425, it holds on the **first** pass: a 38,808-token prompt read in 101 s with steady progress,
no watchdog trip.

This is the Linux/CUDA/stock-engine counterpart to #1425 (Windows, NVFP4 fork: `direct` ≈ 34 MB/s), and the
case where `--ple-io ram` (#746, #605) is **not available** because the host cannot spare 28.8 GB.

## Environment

- **Engine:** stock `0.1.40.2` (`strata serve: stall report (engine 0.1.40.2)`), served via the pack's own
  `serve/server.py --engine strata`; backend cuda.
- **GPU:** RTX 3090 24 GB (1×), `--vram-reserve-mib 700`.
- **Host:** Ubuntu, 62 GiB RAM, 9 GiB swap (≈5 GiB already used); ~47 GiB resident for the engine.
- **Disk:** WD Blue SN580 1 TB NVMe — `/sys/block/nvme0n1/queue/rotational` = `0`, **DRAM-less** (no host
  DRAM for the FTL map).
- **Model:** `Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS` (two shards) + `--ple-gguf` shard **28.8 GB**,
  `--expert-profile … --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp … --max-context 131072
  --kv int8 --vision`. Same result observed with the IQ3_S pack.
- **`--ple-io`:** default (`direct`) unless stated; `--ple-row-cache` / `--ple-inflight` left at defaults
  (1 048 576 / 256) in both arms.

## Reproduction

Fresh (diverse) prompt of agent-session size through the OpenAI-compatible endpoint, on a freshly started
engine so nothing of that prompt is cached:

- `--ple-io direct`: engine log goes to
  ```
  strata serve: no progress for 60 s during a request (reading the prompt (batched): layer 1 of the prompt chunk from token 0) - stopping the engine so the server starts it again (issue #29)
  strata serve: stall report (engine 0.1.40.2): stage "reading the prompt (batched): layer 1 of the prompt chunk from token 0" for 60 s; 0 layers served since the last finished step
    threads waiting on the disk (state D): 2 of 13 (rq_qos_wait x2) - the engine is waiting for the drive: a rotational or failing disk with --ple-io direct (#605: --ple-io ram), or a drive too slow for the reads asked of it
    memory: 47676 MiB resident, 1485 MiB in swap, 13493 MiB RAM available; 305019 major page faults so far
  ```
  Reproduced on every attempt with an 18 672-token prompt and a 30 731-token prompt; small prompts
  (< 200 tokens) finish but are also slow to read.
- `--ple-io mmap`: same prompts, same engine, one flag changed — all pass (numbers below).

## Measurements

Cold prompt reads (no prefix reuse), from the engine's own `prompt … read in … ms` lines:

| prompt | `--ple-io direct` | `--ple-io mmap` |
|---|---|---|
| 59 tokens | 16 145 ms (3.7 tok/s) | — |
| 96–170 tokens | 11 838 / 20 374 ms (7.3 / 8.3 tok/s) | — |
| 18 672 tokens | never posted (deadline hit at layer 1) | **36 780 ms (507.7 tok/s)** |
| 30 731 tokens | never posted | **104 818 ms (293.2 tok/s)** |
| 38 808 tokens | not attempted | **101 255 ms (383.3 tok/s)** |

End-to-end through the client path (same prompt size, both arms):

- `direct`: request fails — `the engine stopped unexpectedly (exit code -6)`; server restarts the engine
  (`experts loaded: 39.97 GiB at 0.83 GiB/s (98 s)`), ~3 min before it can serve again.
- `mmap`: 18 672-token agent prompt answers in 73 s wall clock (prompt read 36.8 s + 486 tokens decoded at
  48.7 tok/s); 30 731-token prompt completes with exit code 0. Prefix reuse stays cheap afterwards
  (`38803 reused + 5 read in 84 ms`).

## Why `--ple-io ram` is not an option here

`--ple-io ram` needs the PLE shard resident (**28.8 GB**), but the host has 62 GiB total with ~47 GiB already
resident for the engine (experts in RAM) and ~13 GiB free — it would push the machine into swap (swap was
already 5 GiB used during the failing run). `mmap` works because the touched row set of a prompt is far
smaller than the whole table and the page cache keeps it across requests, so it costs little RAM.

## The stall-report hint is misleading for this class of disk

`a rotational or failing disk with --ple-io direct` sent me looking at the drive first. A DRAM-less NVMe is
neither rotational nor failing — it does 1.4 GB/s sequentially — but random 4 K reads of ~90-byte rows pay an
extra FTL metadata read each, which is exactly what `direct` (O_DIRECT, no OS cache) does on the prompt path.
Suggest mentioning "DRAM-less SSD / random small reads" (and `--ple-io mmap`) in that hint.

## What would help

1. Should `--ple-io mmap` be the documented recommendation for DRAM-less NVMe when `ram` is unaffordable?
   (It is the only setting that works here, and it holds on the cold pass — unlike the 156 MB/s cold pass
   reported in #1425.)
2. The 60 s watchdog has no slack for a *slow but progressing* read: on `direct` it fires at layer 1 with
   nothing served. Is making it configurable, or exempting the prompt-read stage, on the table?
3. Would raising `--ple-row-cache` / `--ple-inflight` (defaults 1 048 576 / 256) help in this regime? I have
   not tried them; the batched/read-ahead PLE read in the style of #1323 sounds like the real fix.
4. Is `--ple-io ram` meant to be the answer whenever RAM allows (per #746/#605) and `mmap` the fallback, or is
   there a third mode planned (e.g. buffered reads with read-ahead) for hosts that can afford neither the
   RAM nor the O_DIRECT penalty?

Related: #1425 (Windows/NVFP4 fork, `direct` ≈ 34 MB/s), #1211 (`ram` null result on fast NVMe),
#605 / #746 (`ram` as the storage lever), #1323 (batched unbuffered reads), #29 (the watchdog).

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.