Pull requests / #374
#374 Prompt path: the first chunk's n-gram (PLE) rows read beside layer 0, 256 at a time (same output)
closed · @sergqwer · 0 コメント · GitHub で見る
Setup & installNVIDIA / CUDAModels & quantsWindows
本文
Layer 1 adds the per-layer embedding (the n-gram table) to every token of a chunk, so a chunk's rows must be read before layer 1. - **Later chunks were fine:** their rows were already read on a thread during the chunk before. - **The first chunk's were read synchronously before layer 0:** the GPU idled while the SSD read them. An nsys trace of a single-chunk 32K prompt shows ~245 ms of it. What changes: - **The first chunk's gather runs on a thread too**, beside the embedding and layer 0. It lands just before layer 1 (`ple_land`: wait, upload, start the next chunk's gather), and a stage that ends before layer 1 lands it after its last layer. - **`--ple-inflight` defaults to 256 (was 64).** - A 32K prompt is 512,720 rows in 213,358 4 KiB reads. At 64 outstanding reads the drive (Samsung 9100 PRO) ran ~620K IOPS with the GPU waiting. At 256 it saturates, and 1024 is no faster. - Decode reads 16 rows a token and is unchanged. The same rows in the same order give the same results. ## Same output First-token logits are bit-identical to 0.1.31 at 8K and 32K. They were dumped with #276's `STRATA_DUMP_FIRST_LOGITS` in test builds of both, with a fixed `--expert-cache 14900`. ## Speed Setup: IQ2_XS (ISTA), RTX 5090, 9950X3D, 128 GB, Windows 11, `--max-context 262144 --kv int8 --prefill auto` (8192-token chunks), `--vram-reserve-mib 1500`. The runs alternated between 0.1.31 and this PR. | | 0.1.31 | this PR | | --- | ---: | ---: | | 32K prompt (cache auto) | 5,986 / 6,013 ms | **5,843 / 5,808 ms** (-2.9%) | | 8K prompt (fixed cache) | 1,714 ms | **1,624 ms** (-5.3%) | | host time blocked on the rows, 32K | 493-503 ms | 411-432 ms | The gain is the first chunk's read, so it is larger with fewer chunks. On our NVFP4 fork, a 32K prompt in one 32K chunk (#282's auto) was 7-9% faster from the overlap and 2% more from the read depth. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
関連リンク
インストール・モデル・リリースへの站内リンク。