Pull requests / #1656
#1656 pipeline-windows: three or more layer-split stages, +14% decode at 4K / +27% at 32K on 4x RX 7900 XT (opt-in); HIP draft-layer scratch race
open · @Cass67 · 0 commentaires · Sur GitHub
BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Description
## Title
Issue: Related: #1154 (CC-David-CC, open): three P4s, three stages, two speculative windows. Searched 2026-10-09 for `pipeline-windows`, `gr_read` and `ss.block.gr`; no issue for the HIP race below.
## Summary
**Why it matters:** `--pipeline-windows 2` is off on three or more stages ("it needs a layer split into exactly two stages"), so a model that needs four cards decodes with one card busy at a time. On 4x RX 7900 XT with IQ3_S a decode window is four stages of 9.9 to 13.1 ms of GPU work run one after another; the hand-off is 2.2-2.6 ms of host staging (`STRATA_DECODE_TIMING`, `STRATA_VERIFY_PROFILE`, `STRATA_SPLIT_TIMING`).
| 4x RX 7900 XT, Flash-Next GSQ-RCO IQ3_S, split 12,24,36, 262K, int8 KV, every expert in VRAM, greedy, reasoning off, 256 tokens | `--pipeline-windows 0` | `--pipeline-windows 2` |
|---|---|---|
| decode at 4,096 prompt tokens (median of 3) | 59.4 tok/s | 67.6 tok/s (67.4-75.6) |
| decode at 32,768 prompt tokens | 51.8 tok/s | 65.6 tok/s |
| ms per window | 44.0-45.2 | 33.8-35.4 (fresh 42.1-43.7, guessed and kept 14.4-15.1) |
| guessed windows kept | - | 29-34 of 68-87 per request |
Every stage but the last is a front stage. The guessed window follows the verified one through them one card behind; only the last stage waits for the verdict, as with two stages. Default off.
HIP fix, independent of the stage count: the draft layer's per-row `gr_read` (the HIP branch; on CUDA the multi-row read uses the draft layer's own buffers) used `ss.block.gr` of the last stage's session, which that stage's verifier uses in `lm_head_mix`. After a kept guess the chain launched at the verdict and the next window's last stage run on the card at once, and the window's head read a clobbered mix.
## What changed
- `src/program/generate.cpp`: the gate takes two or more stages (`--pipeline-windows 1` and `--adapt-async 1` stay two-stage and are refused on more, with a log line). A stream, an odd verifier and an odd hand-off for every stage and boundary; a pair of GDN snapshots with side stream and events per front stage. The decode loop keeps per window its front progress, an in-flight flag and the front stages it was committed on: B's speculative commit-and-launch runs per front stage; at the verdict A is committed on the front stages B never reached (A.T when the guess is kept, a + 1 otherwise); the undo restores and replays per committed stage. Drain, fences and stats cover every stage. The pipelined prompt read stays two-stage.
- `src/core/mtp.cpp`, `include/strata/core/mtp.hpp`: the draft layer carves its own `GrWorkspace` for that `gr_read`.
- `docs/MULTI_GPU.md`: two or more stages, cost, the off list, the measurements, the HIP note.
- Debug (`STRATA_PIPELINE_DEBUG=1`): `STRATA_PIPELINE_SPEC_DEPTH=k` keeps B to the first k front stages; the `STRATA_PIPELINE_LOG` verdict line also prints B's front progress, A's committed stages and the emitted tokens.
## Extra Notes
Measured on v0.1.41 (fb58e0d) plus this change; the branch is cut from main fb58e0d. 4x RX 7900 XT (gfx1100), i9-9900K, 60.4 GiB RAM, ROCm 10.0.0, Linux; `STRATA_SPLIT_OWN=1 STRATA_STAGE_TRIM=1 STRATA_PF_FUSED=1 STRATA_PF_GEMM=1 STRATA_PF_SWITCH_MIN_T=4096`. Same binary in every arm, only the flag differs; each arm a fresh process; prompts from repository text sized with the pack's tokenizer, unique first line, cache_n 0 on all. The engine's own `prompt_per_second` / `predicted_per_second` and `strata pipeline` lines.
Text equality (`STRATA_IQ_MT_MIN=1 --pcie-frac 0 --adapt-every 0`, greedy, 4 prompts x 256 tokens, 100% of experts in VRAM in every arm): identical to the serial loop, 4 of 4. Also 4 of 4 with `STRATA_PIPELINE_FORCE_MISS=1` (every guess rolled back), `STRATA_PIPELINE_THETA=2` (no guesses) and `STRATA_PIPELINE_SPEC_DEPTH` 1 and 2. This build's serial loop matched v0.1.41's text, 4 of 4.
HIP race: before the draft-layer scratch, full depth was 0 of 4 while depth 1 and 2 were 4 of 4. Each first divergence was in the window after a kept guess whose B had finished every front stage before the verdict, so its last stage launched while the chain ran. Holding that launch until the chain ended (a debug gate, not in this change) gave 4 of 4, and so does the draft layer's own scratch. On two stages the same overlap follows a kept guess whenever B has finished stage 0; not reproduced here, this box needs four cards for the model.
Against #1154: there, three stages with two speculative windows on CUDA (Pascal) and new commit-limit and VMM code; here, any number of stages with one speculative window one card behind, the two-stage machinery extended, on HIP. Both edit the same `generate.cpp` sites.
Not tested: a two-stage split after this change, three stages, CUDA, SYCL, sampled decoding, `--suffix-draft`, images. The snapshot pairs on later front stages are allocated after the expert caches and not kept out of them.
Checked: HIP Release build of `strata` (gfx1100, ROCm 10.0.0). ctest and the Python tests were not run.
Sur le site
Liens install, modèles, releases.