Pull requests / #1516
#1516 PLE: a prompt chunk's rows through 8 batch readers (a 32K chunk -3%, the same rows)
open · @sergqwer · 0 commentaires · Sur GitHub
Description
With `--ple-io direct` (the default) a prompt chunk's PLE rows are one batched request to `PleReader`: one worker thread submits the page reads and reaps every completion. A 32K prompt reads ~213K pages of the n-gram table. In 8192-token chunks (the default `--prefill auto`) that read hides behind layer 0 and the chunk before. In one 32768-token chunk (`--prefill auto:32768`, #282) it does not, and layer 1 waits for it: ``` IQ2_XS, 32K prompt, --prefill auto:32768, RTX 5090 + Samsung 9100 PRO strata generate: prefill 32036 tokens in 1 chunks, 4095.7 ms (...); PLE 161.3 ms ple io: 512720 rows, 0.0% row-cache hits, 211871 SSD reads (886.6 MB), read p50 278 us p99 776 us, blocked 264.030 ms total (submit 25.744 ms) ``` The SSD is not the limit. One thread reaping every completion is. On top of that come the job building before the first read and 8 KiB reads for rows that cross a page boundary. **Change:** `PleReader::read_batch` serves a prompt chunk's rows (256+ tokens, Direct mode only). - Page jobs: one 4 KiB read per page. A row across a page boundary is a use of both pages, each copying its part. - The jobs are deduplicated through a compact table and read in the order of the first row that needs them. - 8 reader threads (`STRATA_PLE_READERS`, 0 = the old path; default 0 on POSIX) each keep 32 reads in flight on their own handle and completion port (`DirectFile` gets an `inline_submit` mode). - `gather_batch` decodes the rows that have landed while the rest are read. - The rows, their bytes and the decode are unchanged. The row cache gets the same rows; only their insertion order differs. - `STRATA_PLE_TRACE=1` prints, per chunk, when its read started, when layer 1 asked for the rows, and how long it waited. - `ple_reader_test --selftest`: `check_batch` with 3 / 1 / 0 readers, rows across pages, duplicates, the table's edges, and a repeat served from the row cache. It passes with row sizes 90 and 110. **Identity** (first-token logits, sha256, `--expert-cache 12000 --pcie-frac 0.25`, d5ea7133 vs this branch): | prompt | d5ea7133 | this PR | |---|---|---| | 2K | a40b6fea84356938 | a40b6fea84356938 | | 8K | d349749a07af4cc5 | d349749a07af4cc5 | | 32K | 050aa170500d5340 | 050aa170500d5340 | | 32K, `--prefill auto:32768` | 4e9b304508ad133c | 4e9b304508ad133c | Every A/B run below had the same logits within its prompt as well. **A/B** on current main: RTX 5090, Ryzen 9 9950X3D, Samsung 9100 PRO, IQ2_XS, `--expert-cache 12000 --pcie-frac 0.25 --vram-reserve-mib 1500`, interleaved pairs, order alternating. The prompt times are from the engine's `prefill N tokens ... ms` line. "Wait at layer 1" is the `PLE` figure of the same line. | prompt | chunks | pairs | d5ea7133 | this PR | paired diff | wait at layer 1 | |---|---|---|---|---|---|---| | 32K | 1 x 32768 (`auto:32768`) | 4 | 4100.6 ± 6.8 ms | 3975.2 ± 0.8 ms | **-125.4 ± 6.6 ms (-3.1%)** | 175.6 -> 48.0 ms | | 32K | 4 x 8192 (`auto`) | 5 | 4841.1 ± 7.1 ms | 4853.5 ± 19.6 ms | +12.4 ± 14.7 ms (noise) | 20.7 -> 20.7 ms | | 8K | 1 x 8034 (`auto`) | 5 | 1320.7 ± 2.8 ms | 1322.0 ± 6.5 ms | +1.4 ± 8.0 ms (noise) | 9.6 -> 9.4 ms | With 8192-token chunks the trace shows 0.0 ms waited at every chunk: the read was already done. So the default gains nothing on this machine. The gain is for prompts read in large chunks: `auto:32768` / `auto:16384`, or a GPU whose layer 0 finishes before the read does. **Not measured:** `auto:16384`; Linux (the readers are off there by default: `DirectFile` on POSIX is a pool of blocking `pread`s, and an inline reader would be one read at a time); SATA or slower NVMe drives; 96K prompts on IQ2_XS. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
Sur le site
Liens install, modèles, releases.