Pull requests / #374

#374 Prompt path: the first chunk's n-gram (PLE) rows read beside layer 0, 256 at a time (same output)

closed · @sergqwer · 0 评论 · 在 GitHub 查看

Setup & installNVIDIA / CUDAModels & quantsWindows

描述

Layer 1 adds the per-layer embedding (the n-gram table) to every token of a chunk, so a chunk's rows must be read before layer 1.
- **Later chunks were fine:** their rows were already read on a thread during the chunk before.
- **The first chunk's were read synchronously before layer 0:** the GPU idled while the SSD read them. An nsys trace of a single-chunk 32K prompt shows ~245 ms of it.

What changes:
- **The first chunk's gather runs on a thread too**, beside the embedding and layer 0. It lands just before layer 1 (`ple_land`: wait, upload, start the next chunk's gather), and a stage that ends before layer 1 lands it after its last layer.
- **`--ple-inflight` defaults to 256 (was 64).**
  - A 32K prompt is 512,720 rows in 213,358 4 KiB reads. At 64 outstanding reads the drive (Samsung 9100 PRO) ran ~620K IOPS with the GPU waiting. At 256 it saturates, and 1024 is no faster.
  - Decode reads 16 rows a token and is unchanged.

The same rows in the same order give the same results.

## Same output

First-token logits are bit-identical to 0.1.31 at 8K and 32K. They were dumped with #276's `STRATA_DUMP_FIRST_LOGITS` in test builds of both, with a fixed `--expert-cache 14900`.

## Speed

Setup: IQ2_XS (ISTA), RTX 5090, 9950X3D, 128 GB, Windows 11, `--max-context 262144 --kv int8 --prefill auto` (8192-token chunks), `--vram-reserve-mib 1500`. The runs alternated between 0.1.31 and this PR.

| | 0.1.31 | this PR |
| --- | ---: | ---: |
| 32K prompt (cache auto) | 5,986 / 6,013 ms | **5,843 / 5,808 ms** (-2.9%) |
| 8K prompt (fixed cache) | 1,714 ms | **1,624 ms** (-5.3%) |
| host time blocked on the rows, 32K | 493-503 ms | 411-432 ms |

The gain is the first chunk's read, so it is larger with fewer chunks. On our NVFP4 fork, a 32K prompt in one 32K chunk (#282's auto) was 7-9% faster from the overlap and 2% more from the read depth.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。