Pull requests / #598

#598 refill: improve multi-GPU processing speed ~27% PP improvement on 4× RTX 3060

closed · @hash13 · 0 评论 · 在 GitHub 查看

Setup & installMulti-GPUNVIDIA / CUDA

描述

## Prefill: allow deeper overlap across multi-GPU pipeline stages

### Problem

In the current multi-GPU prefill path, every stage launches the next stage asynchronously, but still waits for that stage to finish before returning.

With more than two GPUs this creates a recursive wait chain:

```text
GPU0 waits for GPU1
GPU1 waits for GPU2
GPU2 waits for GPU3
```

As a result, intermediate stages cannot accept the next prompt chunk while later GPUs are still processing the previous one.

On a 4-GPU setup this can lead to visibly staggered GPU utilization during prompt processing instead of a fully filled pipeline.

### Change

This PR changes the lifetime of the per-stage asynchronous hand-off.

Intermediate stages now return after handing their chunk to the direct successor instead of draining the complete remaining GPU chain.

The pipeline is drained only once, when the outer/root `Prefill::run()` has submitted all prompt chunks.

Conceptually, the current behavior is closer to:

```text
GPU0: C0 C1 ---- C2 ----
GPU1:    C0 ---- C1 ----
GPU2:       C0 ---- C1
GPU3:          C0 ----
```

The intended behavior after this change is:

```text
GPU0: C0 C1 C2 C3 C4
GPU1:    C0 C1 C2 C3
GPU2:       C0 C1 C2
GPU3:          C0 C1
```

This does not introduce tensor parallelism and does not change the CUDA kernels or model computation.

It only increases overlap between existing layer-split prefill stages.

### Implementation

The change:

- moves the successor `std::future` from local `Prefill::run()` state to the `Prefill` instance
- keeps one outstanding job per direct successor
- prevents intermediate stages from recursively draining all later stages
- preserves the existing double-buffered host hand-off buffers
- drains the complete pipeline once at the end of the root prefill call
- keeps each individual GPU stage serial; two chunks are never executed concurrently on the same `Prefill` instance

### Expected effect

The main benefit should be visible on configurations with 3 or more GPUs and prompts large enough to contain several prefill chunks.

For example, with four GPUs and a 2K prefill chunk size, a 32K prompt provides enough chunks to keep all stages active concurrently after pipeline warm-up.

Short prompts or a single prefill chunk should see little or no improvement.

Decode/token-generation performance is not expected to change.

### Motivation / observed behavior

On a 4× RTX 3060 system, prompt processing currently shows GPU activity progressing largely stage by stage, while decode performance is normal.

This suggests that part of the low prefill throughput may come from the recursive synchronization between layer-split stages rather than raw GPU compute.

This PR is intended as a small, isolated change to test that hypothesis without changing expert streaming, tensor parallelism, CUDA kernels, or the model architecture.

The most useful validation is whether multiple GPUs stay active concurrently during long prompt prefill and whether prompt throughput improves without affecting model output or decode performance.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。