Pull requests / #936
#936 layer split: overlap the stages' hand-off (STRATA_SPLIT_OVERLAP=1, op…
closed · @bsvinay · 0 comentarios · En GitHub
BenchmarksMulti-GPUAMD / HIPModels & quantsDocumentationWindowsLinux
Descripción
On 2x RX 7900 XTX the layer split decoded slower than one card (53/45 tok/s vs ~70/68, see #642). The cause is the order inside every verify window: stage 1 finishes, the host syncs its stream, then stages and launches stage 2's graph. Under HIP both the sync and the graph launch are expensive, and they sit on the critical path of every window while both cards are idle. With STRATA_SPLIT_OVERLAP=1 (off by default; stages on different devices only): - each hand-off buffer gets a ready word (mapped pinned memory); - the writing stage's graph ends with `handoff_publish`: the hand-off payload with volatile stores (RDNA keeps plain stores to mapped memory in L2 - the same reason as doorbell_ring_kernel's volatile store), a system fence, and the last block raises the flag; - the reading stage's graph starts with `wait_flag_ge(flag, 1)`; - `Verifier::run` captures the next stage's window on its own thread, then stages and launches it on a helper thread (`prelaunch`) while it serves its own layers, goes straight on to serving the next stage's layers, and syncs the earlier stage only once the chain is done; - `release_gpu_waits` also raises the hand-in flag, and an early return after the prelaunch releases the next stage (#267). Batch windows, --split-device 0 and the default path are unchanged. Docs: a section in docs/MULTI_GPU.md. Measured on 2x RX 7900 XTX (gfx1100, ROCm, Linux), Ryzen 9 9900X, 52 GB RAM, Swift 1.5 IQ3_XXS, 128K context, 8-bit KV, split layers 0-25 / 26-47; 512-token greedy answers, median of 8 runs: | | serial (default) | STRATA_SPLIT_OVERLAP=1 | one card | |----------------------|------------------|------------------------|----------| | decode, code prompt | 53 tok/s | 70.1 tok/s | ~70 | | decode, prose prompt | 45 tok/s | 55.9 tok/s | ~68 | - Greedy output byte-identical with the flag at 0 and 1 (same prompt, both runs checked in the engine log: overlapped / serial). - Setting it back to 0 returns the old speeds (53/45). - Quiz set: 47/51 on the split, same as one card. - Prefill unchanged (1,232 / 1,763 / 2,147 tok/s at 4K / 16K / 32K). Not measured on NVIDIA, where a sync and a graph launch cost less - that is why it is opt-in. Happy to change the default or the switch name as you prefer.
En el sitio
Enlaces a install, modelos, releases.