Pull requests / #663

#663 prefill: a layer split's idle card streams a share of a one-chunk prompt's experts

closed · @anon761 · 0 comentários · No GitHub

BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

Descrição

## What

A prompt that fits one chunk runs a layer split's stages one after the other: while one card reads its layers, the other idles. `--peer-device` already has the mechanism to use such a card - `PeerPrefill`'s peer streaming (a share of the experts the primary streams goes over the peer's own PCIe link and is computed there) - but it needs P2P and is refused with `--layer-split`. Most consumer pairs have no P2P (two RTX 3090 here: `nvidia-smi topo -p2p r` says `CNS`).

This makes `PeerPrefill` usable as a **stream-only peer between the stages**:

- `Prefill::set_stage_helper`: each stage's prompt path gets the next stage's (the last one the first's). A stream-only `PeerPrefill` holds no experts (`peer == nullptr`); every use of the peer tier is guarded.
- `bind_stage_helper`, per prompt: points it at the idle stage's own lent prompt buffers (`mixed`, `Xq`, `GU`, `H`, `Hq`, `Dm`, its ring, its MMQ context). **No VRAM of its own**; `serve`'s `lend()` lays every stage's buffers out at once before the prompt.
- **Without P2P** the activations, the row table, the group bounds and the result rows go through mapped host memory and are read by copy kernels on the device that needs them (a new `copy_f32_wide`, 16-byte loads). My first version used copy-engine uploads; they queued behind the ring's expert blobs - the same reason the grouping tables live in mapped memory - and cost ~20 ms per layer, which ate the whole gain (the `combine` phase went 24 -> 407 ms). `cudaMemcpyPeerAsync` without P2P is staged by the driver.
- Only for one-chunk prompts on the MMQ path (not the fused prompt kernels, not with `--peer-device`), with a share that falls with the chunk - `min(0.45, 0.5 - T/16384)`, off below 0.3 (~3.3K tokens) - from the sweep below.
- `STRATA_PREFILL_HELP=0` turns it off, `STRATA_PREFILL_HELP_FRAC=f` fixes the share. `docs/MULTI_GPU.md` describes it.

The `--peer-device` path is unchanged (its `ctx` becomes `run_ctx`, a pointer, so the helper can run in the idle stage's context).

## Measured

2x RTX 3090 (PCIe 4.0 x16 each, no P2P), EPYC 7413, Unsloth UD-Q4_K_XL, layer split auto, 262K context, `--kv int8 --pcie-frac 0 --ple-io ram --adapt-swaps 0`, synthetic prompts (a repeated paragraph, unique per run) over `/v1/chat/completions`, 3 runs each, same binary:

| prompt | `STRATA_PREFILL_HELP=0` | default (help) | |
| --- | ---: | ---: | ---: |
| 1.1K | 496 tok/s | 723 tok/s | +46% |
| 1.5K | 674 | 833 | +23% |
| 2K | 878 | 1,128 | +28% |
| 2.5K | 1,017 | 1,296 | +27% |
| 3K | 1,260 | 1,440 | +14% |
| 4K | 1,605 | 1,607 | (off by the rule) |
| 8K | 2,215 | 2,221 | (off by the rule) |

Decode unchanged. The share sweep behind the rule (`STRATA_PREFILL_HELP_FRAC` 0.2 / 0.3 / 0.4 against none): 1.5K 749 / 820 / 908 (675), 3K 1,335 / 1,437 / 1,328 (1,262), 4K 1,622 / 1,620 / 1,440 (1,609), 6K 1,811 / 1,660 / 1,471 (2,038), 8K 1,895 / 1,804 / 1,663 (2,230).

## Output

Repeatable: the same prompt gives the same answer, run after run (`--adapt-swaps 0`, four fixed prompts of 1.2-3.1K tokens, twice). It is **not bit-identical** to `STRATA_PREFILL_HELP=0`'s: the helped rows are computed in another MMQ grouping and round differently - like `STRATA_SPLIT_OWN`. I made it default-on because the default (adaptive swaps on) is not bit-repeatable across requests anyway and the gain is large for the short prompts agents send; if you prefer it opt-in like `STRATA_SPLIT_OWN`, that is a one-line change and I'm happy to flip it.

Not tested: a P2P pair (the helper always takes the mapped route), HIP, three or more stages (the code takes the next stage's card; only two were available).

No site

Links install, modelos, releases.