Pull requests / #1120

#1120 layer split: overlap the stages' hand-off (STRATA_SPLIT_OVERLAP=1, opt-in)

open · @bsvinay · 0 comentários · No GitHub

BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

Descrição

Resubmitted from #936, which GitHub closed when main's history was rewritten. Same 4 commits, cherry-picked onto the new main (no conflicts). The review discussion and test results from others are in #936.

On 2x RX 7900 XTX the layer split decoded slower than one card. The cause is the order inside every verify window: stage 1 finishes, the host syncs its stream, then stages and launches stage 2's graph. Under HIP both the sync and the graph launch are expensive, and they sit on the critical path of every window while both cards are idle.

With STRATA_SPLIT_OVERLAP=1 (off by default; stages on different devices only):
- each hand-off buffer gets a ready word (mapped pinned memory);
- the writing stage's graph ends with `handoff_publish`: the payload with volatile stores (RDNA keeps plain stores to mapped memory in L2, same reason as doorbell_ring_kernel), a system fence, and the last block raises the flag;
- the reading stage's graph starts with a wait on that flag;
- `Verifier::run` launches the next stage's window on a helper thread while it serves its own layers, and syncs the earlier stage only once the chain is done;
- `release_gpu_waits` also raises the hand-in flag, and an early return after the early launch releases the next stage (#267).

Changes since #936 (from the review there):
- the GPU-side hand-off wait is bounded: after STRATA_SPLIT_WAIT_MS (default 30000, 0 = no limit) it sets an error word and lets the graph finish; the host then fails that window with a clear error instead of using a stale hand-off;
- a test hook, STRATA_TEST_HANDOFF_DROP=N, leaves the flag down in the Nth window (the publish reads it only when the hook is set);
- the early launch keeps the #871 graph variant the next stage picked;
- the overlap switches itself off with --pipeline-windows (#859), which overlaps the stages its own way.

Measured on 2x RX 7900 XTX (gfx1100, ROCm 10.0.0, Linux), Ryzen 9 9900X, 52 GB RAM, Swift 1.5 IQ3_XXS, 128K context, 8-bit KV, split 0-25 / 26-47, 512-token greedy answers:

|                      | serial (default) | STRATA_SPLIT_OVERLAP=1 |
|----------------------|------------------|------------------------|
| decode, code prompt  | 53 tok/s         | 69.6 tok/s             |
| decode, prose prompt | 45 tok/s         | 57.8 tok/s             |

- Greedy output byte-identical with the flag at 0 and 1.
- Fault test with the hook: with STRATA_SPLIT_WAIT_MS=2000 the request fails after ~2 s with the hand-off error (HTTP 400) and both GPUs are idle; with the default, the #267 host timeout fires at 20 s, the GPUs are idle and the server restarts the engine (503); with the overlap off the hook does nothing.
- From #936: +1.6% decode on 4x V100 (@justxiami), no change on 4x / 2x RTX 5080 (@blange48) and on RDNA4 (@aswin-dot-R, byte-identical output). So it mostly helps RDNA3 - that is why it stays opt-in.

No site

Links install, modelos, releases.