Pull requests / #1516

#1516 PLE: a prompt chunk's rows through 8 batch readers (a 32K chunk -3%, the same rows)

open · @sergqwer · 0 commentaires · Sur GitHub

NVIDIA / CUDALinux

Description

With `--ple-io direct` (the default) a prompt chunk's PLE rows are one batched request to `PleReader`: one worker thread submits the page reads and reaps every completion. A 32K prompt reads ~213K pages of the n-gram table. In 8192-token chunks (the default `--prefill auto`) that read hides behind layer 0 and the chunk before. In one 32768-token chunk (`--prefill auto:32768`, #282) it does not, and layer 1 waits for it:

```
IQ2_XS, 32K prompt, --prefill auto:32768, RTX 5090 + Samsung 9100 PRO
strata generate: prefill 32036 tokens in 1 chunks, 4095.7 ms (...); PLE 161.3 ms
  ple io: 512720 rows, 0.0% row-cache hits, 211871 SSD reads (886.6 MB), read p50 278 us p99 776 us, blocked 264.030 ms total (submit 25.744 ms)
```

The SSD is not the limit. One thread reaping every completion is. On top of that come the job building before the first read and 8 KiB reads for rows that cross a page boundary.

**Change:** `PleReader::read_batch` serves a prompt chunk's rows (256+ tokens, Direct mode only).

- Page jobs: one 4 KiB read per page. A row across a page boundary is a use of both pages, each copying its part.
- The jobs are deduplicated through a compact table and read in the order of the first row that needs them.
- 8 reader threads (`STRATA_PLE_READERS`, 0 = the old path; default 0 on POSIX) each keep 32 reads in flight on their own handle and completion port (`DirectFile` gets an `inline_submit` mode).
- `gather_batch` decodes the rows that have landed while the rest are read.
- The rows, their bytes and the decode are unchanged. The row cache gets the same rows; only their insertion order differs.
- `STRATA_PLE_TRACE=1` prints, per chunk, when its read started, when layer 1 asked for the rows, and how long it waited.
- `ple_reader_test --selftest`: `check_batch` with 3 / 1 / 0 readers, rows across pages, duplicates, the table's edges, and a repeat served from the row cache. It passes with row sizes 90 and 110.

**Identity** (first-token logits, sha256, `--expert-cache 12000 --pcie-frac 0.25`, d5ea7133 vs this branch):

| prompt | d5ea7133 | this PR |
|---|---|---|
| 2K | a40b6fea84356938 | a40b6fea84356938 |
| 8K | d349749a07af4cc5 | d349749a07af4cc5 |
| 32K | 050aa170500d5340 | 050aa170500d5340 |
| 32K, `--prefill auto:32768` | 4e9b304508ad133c | 4e9b304508ad133c |

Every A/B run below had the same logits within its prompt as well.

**A/B** on current main: RTX 5090, Ryzen 9 9950X3D, Samsung 9100 PRO, IQ2_XS, `--expert-cache 12000 --pcie-frac 0.25 --vram-reserve-mib 1500`, interleaved pairs, order alternating. The prompt times are from the engine's `prefill N tokens ... ms` line. "Wait at layer 1" is the `PLE` figure of the same line.

| prompt | chunks | pairs | d5ea7133 | this PR | paired diff | wait at layer 1 |
|---|---|---|---|---|---|---|
| 32K | 1 x 32768 (`auto:32768`) | 4 | 4100.6 ± 6.8 ms | 3975.2 ± 0.8 ms | **-125.4 ± 6.6 ms (-3.1%)** | 175.6 -> 48.0 ms |
| 32K | 4 x 8192 (`auto`) | 5 | 4841.1 ± 7.1 ms | 4853.5 ± 19.6 ms | +12.4 ± 14.7 ms (noise) | 20.7 -> 20.7 ms |
| 8K | 1 x 8034 (`auto`) | 5 | 1320.7 ± 2.8 ms | 1322.0 ± 6.5 ms | +1.4 ± 8.0 ms (noise) | 9.6 -> 9.4 ms |

With 8192-token chunks the trace shows 0.0 ms waited at every chunk: the read was already done. So the default gains nothing on this machine. The gain is for prompts read in large chunks: `auto:32768` / `auto:16384`, or a GPU whose layer 0 finishes before the read does.

**Not measured:** `auto:16384`; Linux (the readers are off there by default: `DirectFile` on POSIX is a pool of blocking `pread`s, and an inline reader would be one read at a time); SATA or slower NVMe drives; 96K prompts on IQ2_XS.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

Sur le site

Liens install, modèles, releases.