Pull requests / #1441
#1441 Prompt chunks sized for a layer split's pipeline (STRATA_PREFILL_PIPE=1, opt-in)
open · @ruibeikaa · 0 comentarios · En GitHub
BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentation
Descripción
## Title Prompt chunks sized for a layer split's pipeline (`STRATA_PREFILL_PIPE=1`, opt-in) ## Summary `--prefill auto` takes the largest chunk the buffers allow (up to 8192). That is right for one card - a chunk costs a pass over its layers' experts plus a part per token, so few large chunks are cheapest - and wrong for a layer split, which reads a prompt as a pipeline: with few large chunks the later stages wait. A 10.7K-token prompt read as a 7680 and a 3011 chunk over four stages kept the first stage busy for 24 of the 36 s (`STRATA_PREFILL_TIMING`), and the four GPUs averaged 1.2 busy (nvidia-smi at 100 ms). With `STRATA_PREFILL_PIPE=1` the first stage reads each prompt in the chunk that minimizes the pipeline's time: (S - 1 + n/C) stage-chunks of (b + a C) for n tokens over S stages is smallest at **C = sqrt(n (b/a) / (S - 1))**, b being a layer's per-chunk pass and a its per-token part. So the chunk grows with the square root of the prompt: ~1024 at 2-4K tokens, 2048 at 10K, 3328 at 30K on four stages. | Prompt tokens | `--prefill auto` | `--prefill 2048` | `--prefill auto` + `STRATA_PREFILL_PIPE=1` (chunk) | | ---: | ---: | ---: | ---: | | 2,013 | 180.9 | 181.4 | **228.0** (1024) | | 3,865 | 207.6 | 276.5 | **312.4** (1024) | | 10,041 | 261.1 | 433.0 | **432.6** (2048) | | 28,998 | 475.7 | 554.1 | **559.7** (3328) | Prompt tok/s on four GP100 dies (PH402 SKU 200), `--layer-split 12,24,36`, Flash-Next IQ3_XXS, 64K context, fresh prompts, two interleaved rounds per policy (each within 0.2%). Against `auto` +18-66%; against the best fixed chunk I had found by hand (2048) the same or better at every length, +26% at 2K. ## What changed - `src/prefill/prefill.cpp`, `include/strata/prefill/prefill.hpp`: `Prefill::pipeline_chunk(n, cap, stages)` and its use in `run_impl`. b/a is in tokens and set mostly by the model (its experts per token) rather than the card: ~1000 for Flash-Next here (b ~ 80 ms and a ~ 0.08 ms per layer, from `STRATA_PREFILL_TIMING` at two chunk sizes); the default is 1024 and `STRATA_PREFILL_PIPE=<b/a>` sets another. The chunk is rounded up to the 256-token grid, at least 512, at most the buffers' chunk, and evened out over the chunks it takes (no short last chunk). A prompt the rule would read in one chunk keeps today's path (and the stage helper for one-chunk prompts). Only the first stage chooses - a later stage is handed one chunk at a time - and the stage count comes from the `next_` chain, so nothing outside the prompt path changes. One log line says what it chose: `strata prefill: 10041 tokens in 2048-token chunks over 4 stages (STRATA_PREFILL_PIPE)`. - `docs/MULTI_GPU.md`: a section with the model, the table and how to measure b/a on another rig. ## Extra Notes - Opt-in because the chunk geometry changes the rounding (as `STRATA_PREFILL_EQUAL`). Without the variable, or with one stage, nothing changes. - Keep `--prefill auto` with it: the buffers then allow long prompts their larger chunks; a fixed `--prefill N` caps the chunk at N. - Measured only on a four-stage split of one model. On two cards S - 1 = 1 makes the chunks larger (e.g. ~3300 at 10K) and the gain smaller; I could not measure a two-card split here. Results from other rigs would show whether the default b/a holds or wants a measured value. - The first engine start of a session read prompts slower (171 instead of 228 tok/s at 2K; seen twice that day, with other settings too - the model's files not yet in RAM); the table uses the later, warm starts. - Built with CUDA 12.9 / MSVC 2022 for sm_60 (the same diff also on top of #1424). Independent of #1424.
En el sitio
Enlaces a install, modelos, releases.