Pull requests / #1099

#1099 layer split: the hand-off over peer access on NVLink (3.3 -> 48 GB/s, +0.9% e2e)

open · @ATIVX928 · 0 commentaires · Sur GitHub

BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAWindows

Description

> Resubmitted from #1023: the original was auto-closed by the 2026-10-06 history cleanup. Rebased onto the new `main` (82f46a8); build and benchmark re-verified on 2x V100.

## What

The layer-split hand-off goes through pinned host RAM today. This adds the peer-access path (`split_handoff_p2p`), split out of #627:

- verify windows: a peer store inside the captured window graph (a copy kernel on the source card; `cudaMemcpyPeerAsync` is rejected inside a capture, the bench records this)
- prompt chunks: `cudaMemcpyPeerAsync` on the source stage, with the receiving stage's chunk buffers on its own card
- `split_p2p_bench` and the MULTI_GPU.md section

The pinned-RAM path remains the default outside the measured configuration.

## Gating

`STRATA_SPLIT_P2P=0` forces the pinned path, `=1` forces the link, unset means auto only on an sm_70 + sm_70 pair inside the `STRATA_EXPERIMENTAL_SM60` build. HIP always keeps the pinned path. The ready-made engine is unchanged.

## Measured

2x V100-SXM2-16GB (NVLink nv2), CUDA 12.8, `split_p2p_bench`, useful bandwidth per hand-off:

| shape | pinned bounce | p2p memcpy | p2p kernel |
| --- | ---: | ---: | ---: |
| verify 8-token window (400 KB) | 2.99 GB/s | 26.24 GB/s | 22.70 GB/s |
| prompt chunk 2048 (84 MB) | 3.31 GB/s | 48.27 GB/s | 45.55 GB/s |

One hand-off: 0.137 ms -> 0.018 ms for the window; 25.33 ms -> 1.738 ms for the 2048-token chunk. Every path is checked byte-for-byte in the bench before it is timed.

## Tests

- `split_p2p_bench` built and run on 2x V100 (both GPUs), all paths byte-identical.
- The engine target builds clean; without P2P the peer pointer is null and the original pinned branches run.

## 32K end-to-end

Same setup; the same binary with `STRATA_SPLIT_P2P=0` as the control, interleaved (3 rounds each, two pairs): on the rebased branch against the new main, dual-card prefill 2213.9 -> 2233.1 tok/s (+0.9%); on the pre-cleanup main the same A/B read 2190.7 -> 2212.9 (+1.0%). Decode is unchanged (within its +-15% session noise).

Sur le site

Liens install, modèles, releases.