Pull requests / #578

#578 Improve multi-GPU decode performance via optimized expert-helper execution

closed · @zhedang · 0 评论 · 在 GitHub 查看

BenchmarksMulti-GPUNVIDIA / CUDAModels & quants

描述

## Summary

This PR improves multi-GPU decode performance by making the primary GPU and
expert-helper GPUs cooperate more efficiently.

The optimized **v2.2 expert-helper path** delivers substantial decode speedups
over both the previous expert-helper path and the automatic layer-split path.

### Speedup summary

| GPUs / Version | Compared against | Mixed | Code | Math20 |
|---|---|---:|---:|---:|
| RTX 5090 + RTX 4090 / 0.1.38 | Original expert-helper | **+28%** | **+63%** | **+42%** |
| RTX 5090 + RTX 4090 / 0.1.38 | Automatic layer split | **+17%** | **+16%** | **+18%** |
| Dual RTX 4090 / 0.1.37 | Original expert-helper | **+63%** | **+132%** | — |
| Dual RTX 4090 / 0.1.37 | Automatic layer split | **+13%** | **+34%** | — |

> Enabled with `--remote-expert-opt`. Disabled by default.

---

## Benchmarks

### RTX 5090 + RTX 4090 — 0.1.38

Decode throughput, in tokens/s:

| Configuration | Mixed | Code | Math20 |
|---|---:|---:|---:|
| Original expert helper | 135.49 | 152.56 | 158.61 |
| Automatic layer split | 147.41 | 213.89 | 191.06 |
| **Optimized expert helper** | **173.04** | **248.38** | **226.00** |

### Dual RTX 4090 — 0.1.37

Decode throughput, in tokens/s:

| Configuration | Mixed | Code |
|---|---:|---:|
| Original expert helper | 88.07 | 85.15 |
| Automatic layer split | 126.86 | 146.95 |
| **Optimized expert helper** | **143.35** | **197.58** |

---

## What changed

- **Complementary expert caches**  
  Keep the primary and helper GPU expert caches complementary, reducing
  duplicated experts and CPU fallback.

- **Lower inter-GPU transfer volume**  
  Reduce weighted expert outputs on helper GPUs before returning results.

- **Skip unnecessary CPU quantization**  
  Avoid CPU activation quantization when all selected experts are served by GPUs.

- **Automatic helper-cache sizing**  
  Add `auto` sizing based on available VRAM, removing the need to manually tune
  expert counts.

The existing primary/helper GPU layout remains unchanged.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。