Pull requests / #598
#598 refill: improve multi-GPU processing speed ~27% PP improvement on 4× RTX 3060
closed · @hash13 · 0 コメント · GitHub で見る
Setup & installMulti-GPUNVIDIA / CUDA
本文
## Prefill: allow deeper overlap across multi-GPU pipeline stages ### Problem In the current multi-GPU prefill path, every stage launches the next stage asynchronously, but still waits for that stage to finish before returning. With more than two GPUs this creates a recursive wait chain: ```text GPU0 waits for GPU1 GPU1 waits for GPU2 GPU2 waits for GPU3 ``` As a result, intermediate stages cannot accept the next prompt chunk while later GPUs are still processing the previous one. On a 4-GPU setup this can lead to visibly staggered GPU utilization during prompt processing instead of a fully filled pipeline. ### Change This PR changes the lifetime of the per-stage asynchronous hand-off. Intermediate stages now return after handing their chunk to the direct successor instead of draining the complete remaining GPU chain. The pipeline is drained only once, when the outer/root `Prefill::run()` has submitted all prompt chunks. Conceptually, the current behavior is closer to: ```text GPU0: C0 C1 ---- C2 ---- GPU1: C0 ---- C1 ---- GPU2: C0 ---- C1 GPU3: C0 ---- ``` The intended behavior after this change is: ```text GPU0: C0 C1 C2 C3 C4 GPU1: C0 C1 C2 C3 GPU2: C0 C1 C2 GPU3: C0 C1 ``` This does not introduce tensor parallelism and does not change the CUDA kernels or model computation. It only increases overlap between existing layer-split prefill stages. ### Implementation The change: - moves the successor `std::future` from local `Prefill::run()` state to the `Prefill` instance - keeps one outstanding job per direct successor - prevents intermediate stages from recursively draining all later stages - preserves the existing double-buffered host hand-off buffers - drains the complete pipeline once at the end of the root prefill call - keeps each individual GPU stage serial; two chunks are never executed concurrently on the same `Prefill` instance ### Expected effect The main benefit should be visible on configurations with 3 or more GPUs and prompts large enough to contain several prefill chunks. For example, with four GPUs and a 2K prefill chunk size, a 32K prompt provides enough chunks to keep all stages active concurrently after pipeline warm-up. Short prompts or a single prefill chunk should see little or no improvement. Decode/token-generation performance is not expected to change. ### Motivation / observed behavior On a 4× RTX 3060 system, prompt processing currently shows GPU activity progressing largely stage by stage, while decode performance is normal. This suggests that part of the low prefill throughput may come from the recursive synchronization between layer-split stages rather than raw GPU compute. This PR is intended as a small, isolated change to test that hypothesis without changing expert streaming, tensor parallelism, CUDA kernels, or the model architecture. The most useful validation is whether multiple GPUs stay active concurrently during long prompt prefill and whether prompt throughput improves without affecting model output or decode performance.
関連リンク
インストール・モデル・リリースへの站内リンク。