Pull requests / #1441

#1441 Prompt chunks sized for a layer split's pipeline (STRATA_PREFILL_PIPE=1, opt-in)

open · @ruibeikaa · 0 comentários · No GitHub

BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentation

Descrição

## Title
Prompt chunks sized for a layer split's pipeline (`STRATA_PREFILL_PIPE=1`, opt-in)

## Summary
`--prefill auto` takes the largest chunk the buffers allow (up to 8192). That is right for one card - a chunk costs a
pass over its layers' experts plus a part per token, so few large chunks are cheapest - and wrong for a layer split,
which reads a prompt as a pipeline: with few large chunks the later stages wait. A 10.7K-token prompt read as a 7680
and a 3011 chunk over four stages kept the first stage busy for 24 of the 36 s (`STRATA_PREFILL_TIMING`), and the
four GPUs averaged 1.2 busy (nvidia-smi at 100 ms).

With `STRATA_PREFILL_PIPE=1` the first stage reads each prompt in the chunk that minimizes the pipeline's time:
(S - 1 + n/C) stage-chunks of (b + a C) for n tokens over S stages is smallest at **C = sqrt(n (b/a) / (S - 1))**,
b being a layer's per-chunk pass and a its per-token part. So the chunk grows with the square root of the prompt:
~1024 at 2-4K tokens, 2048 at 10K, 3328 at 30K on four stages.

| Prompt tokens | `--prefill auto` | `--prefill 2048` | `--prefill auto` + `STRATA_PREFILL_PIPE=1` (chunk) |
| ---: | ---: | ---: | ---: |
| 2,013 | 180.9 | 181.4 | **228.0** (1024) |
| 3,865 | 207.6 | 276.5 | **312.4** (1024) |
| 10,041 | 261.1 | 433.0 | **432.6** (2048) |
| 28,998 | 475.7 | 554.1 | **559.7** (3328) |

Prompt tok/s on four GP100 dies (PH402 SKU 200), `--layer-split 12,24,36`, Flash-Next IQ3_XXS, 64K context, fresh
prompts, two interleaved rounds per policy (each within 0.2%). Against `auto` +18-66%; against the best fixed chunk
I had found by hand (2048) the same or better at every length, +26% at 2K.

## What changed
- `src/prefill/prefill.cpp`, `include/strata/prefill/prefill.hpp`: `Prefill::pipeline_chunk(n, cap, stages)` and its
  use in `run_impl`. b/a is in tokens and set mostly by the model (its experts per token) rather than the card: ~1000
  for Flash-Next here (b ~ 80 ms and a ~ 0.08 ms per layer, from `STRATA_PREFILL_TIMING` at two chunk sizes); the
  default is 1024 and `STRATA_PREFILL_PIPE=<b/a>` sets another. The chunk is rounded up to the 256-token grid, at
  least 512, at most the buffers' chunk, and evened out over the chunks it takes (no short last chunk). A prompt the
  rule would read in one chunk keeps today's path (and the stage helper for one-chunk prompts). Only the first stage
  chooses - a later stage is handed one chunk at a time - and the stage count comes from the `next_` chain, so nothing
  outside the prompt path changes. One log line says what it chose:
  `strata prefill: 10041 tokens in 2048-token chunks over 4 stages (STRATA_PREFILL_PIPE)`.
- `docs/MULTI_GPU.md`: a section with the model, the table and how to measure b/a on another rig.

## Extra Notes
- Opt-in because the chunk geometry changes the rounding (as `STRATA_PREFILL_EQUAL`). Without the variable, or with
  one stage, nothing changes.
- Keep `--prefill auto` with it: the buffers then allow long prompts their larger chunks; a fixed `--prefill N` caps
  the chunk at N.
- Measured only on a four-stage split of one model. On two cards S - 1 = 1 makes the chunks larger (e.g. ~3300 at
  10K) and the gain smaller; I could not measure a two-card split here. Results from other rigs would show whether
  the default b/a holds or wants a measured value.
- The first engine start of a session read prompts slower (171 instead of 228 tok/s at 2K; seen twice that day, with
  other settings too - the model's files not yet in RAM); the table uses the later, warm starts.
- Built with CUDA 12.9 / MSVC 2022 for sm_60 (the same diff also on top of #1424). Independent of #1424.

No site

Links install, modelos, releases.