Pull requests / #903
#903 prefill: a prompt reads in equal chunks; a short streaming tail stays (#693, rebased)
closed · @Zauberio · 0 commentaires · Sur GitHub
BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsWindows
Description
Rebased variant of @architectds's equal-chunk idea from #693 — the idea and the original data are theirs; the `auto:16384` sweep part of the original is superseded by #583's bisection, so this keeps only the chunk-splitting change. q8atnight asked for a fresh PR against main before adding a 3090 row (#693), so here it is. ## What this does `request_chunk` currently returns `max_chunk` for every full chunk and lets the tail be whatever is left. This splits a segment into **equal chunks** instead: `n = ceil(tokens / max_chunk)` chunks of `ceil(tokens / n)`, rounded to 256 — with one exception kept from the current behavior: when the last piece would fall below `Prefill::stream_all_min_tokens()`, full chunks + short tail stay, because a sub-threshold tail moves only the experts its own tokens route to, while equal chunks would all stream every expert (measured: 16,402 tokens in 3 equal chunks read 22% *slower* than 2 × 8,192 + 18). The cost model is the streaming threshold: a chunk at or above `stream_all_min_tokens()` streams every expert the GPU does not hold (~1.7 s each on PCIe 3.0, whatever its length), so a long prompt's cost is its *number* of such chunks — a bigger chunk pays only where it saves one, and equal chunks borrow no more slots than that count needs. ## Measured (RTX 5070 Ti 16 GB, Windows, IQ3_S, `--prefill auto:32768`, engine built from `main` at `6f32ec0`) Two binaries — stock main (`97E3A38E…`) and this patch (`4088639E…`), same toolchain, same base commit. ABBA swap-boot legs (cross-boot variance here is ~5%, single boots prove nothing), 3 greedy probes per leg, steady-state runs averaged: | prompt | stock layout | equal-chunk layout | stock tok/s | patched tok/s | delta | |---|---|---|---|---|---| | 22,588 | 20,224 + 2,364 | 11,520 + 11,068 | 3,316 | 3,457 | **+4.3%** | | 26,637 | 20,224 + 6,413 | 13,568 + 13,069 | 3,443 | 3,507 | +1.9% | | 40,587 | 20,224 + 20,224 + 139 | *identical* | 3,413 | 3,451 | +1.1% (drift) | The 40K row is a built-in null control: bisection's tail rule keeps the layout identical there, so its +1.1% is pure build-pair drift — call the real gain ≈ +3% at 22.6K, marginal at 26.6K, zero where the layout doesn't change. Consistent with the original 5070 Ti claim (+2.2% at 20,036). `bench/results/2026-10-03-prompt-chunks/` carries the raw ABBA JSON and needle probes. ## For @q8atnight - **2× 5060 Ti rig:** confirmed from our side too — with a pinned `--prefill 6144` this never fires (the rule only re-splits totals above the threshold), so nothing changes there. - **Exactness rule:** acknowledged and disclosed below. The needle-retrieval probes (unique code buried at ~60% depth of a ~12K-token prompt) answer correctly on both builds, but see the caveat. ## Honest caveat: not bit-identical Different chunk boundaries shift accumulation order, so a fixed prompt's greedy continuation **differs between builds** (both coherent, needles retrieve). Anything asserting exact tokens across prefill layouts will need updating — per your gate rule this needs a fresh long-gate reference before adoption, which is exactly what we're not claiming here. ## Safety on constrained boxes This never *raises* a chunk above what bisection picked — it only splits an already-chosen total into equal pieces (22,588: 20,224+2,364 → 11,520+11,068). On layer-split / low-RAM boxes the chunks get smaller, not bigger.
Sur le site
Liens install, modèles, releases.