Pull requests / #1023
#1023 layer split: the hand-off over peer access on NVLink (3.3 -> 48 GB/s, +1.0% e2e)
closed · @ATIVX928 · 0 comentários · No GitHub
BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAWindows
Descrição
## What The layer-split hand-off goes through pinned host RAM today. This adds the peer-access path (`split_handoff_p2p`), split out of #627: - verify windows: a peer store inside the captured window graph (a copy kernel on the source card; `cudaMemcpyPeerAsync` is rejected inside a capture, the bench records this) - prompt chunks: `cudaMemcpyPeerAsync` on the source stage, with the receiving stage's chunk buffers on its own card - `split_p2p_bench` and the MULTI_GPU.md section The pinned-RAM path remains the default outside the measured configuration. ## Gating `STRATA_SPLIT_P2P=0` forces the pinned path, `=1` forces the link, unset means auto only on an sm_70 + sm_70 pair inside the `STRATA_EXPERIMENTAL_SM60` build. HIP always keeps the pinned path. The ready-made engine is unchanged. ## Measured 2x V100-SXM2-16GB (NVLink nv2), CUDA 12.8, `split_p2p_bench`, useful bandwidth per hand-off: | shape | pinned bounce | p2p memcpy | p2p kernel | | --- | ---: | ---: | ---: | | verify 8-token window (400 KB) | 2.99 GB/s | 26.24 GB/s | 22.70 GB/s | | prompt chunk 2048 (84 MB) | 3.31 GB/s | 48.27 GB/s | 45.55 GB/s | One hand-off: 0.137 ms -> 0.018 ms for the window; 25.33 ms -> 1.738 ms for the 2048-token chunk. Every path is checked byte-for-byte in the bench before it is timed. ## Tests - `split_p2p_bench` built and run on 2x V100 (both GPUs), all paths byte-identical. - The engine target builds clean; without P2P the peer pointer is null and the original pinned branches run. ## 32K end-to-end Same setup; the same binary with `STRATA_SPLIT_P2P=0` as the control, interleaved (3 rounds each, two pairs): dual-card prefill 2190.7 -> 2212.9 tok/s (+1.0%). Decode is unchanged (within its +-15% session noise).
No site
Links install, modelos, releases.