Pull requests / #578
#578 Improve multi-GPU decode performance via optimized expert-helper execution
closed · @zhedang · 0 评论 · 在 GitHub 查看
BenchmarksMulti-GPUNVIDIA / CUDAModels & quants
描述
## Summary This PR improves multi-GPU decode performance by making the primary GPU and expert-helper GPUs cooperate more efficiently. The optimized **v2.2 expert-helper path** delivers substantial decode speedups over both the previous expert-helper path and the automatic layer-split path. ### Speedup summary | GPUs / Version | Compared against | Mixed | Code | Math20 | |---|---|---:|---:|---:| | RTX 5090 + RTX 4090 / 0.1.38 | Original expert-helper | **+28%** | **+63%** | **+42%** | | RTX 5090 + RTX 4090 / 0.1.38 | Automatic layer split | **+17%** | **+16%** | **+18%** | | Dual RTX 4090 / 0.1.37 | Original expert-helper | **+63%** | **+132%** | — | | Dual RTX 4090 / 0.1.37 | Automatic layer split | **+13%** | **+34%** | — | > Enabled with `--remote-expert-opt`. Disabled by default. --- ## Benchmarks ### RTX 5090 + RTX 4090 — 0.1.38 Decode throughput, in tokens/s: | Configuration | Mixed | Code | Math20 | |---|---:|---:|---:| | Original expert helper | 135.49 | 152.56 | 158.61 | | Automatic layer split | 147.41 | 213.89 | 191.06 | | **Optimized expert helper** | **173.04** | **248.38** | **226.00** | ### Dual RTX 4090 — 0.1.37 Decode throughput, in tokens/s: | Configuration | Mixed | Code | |---|---:|---:| | Original expert helper | 88.07 | 85.15 | | Automatic layer split | 126.86 | 146.95 | | **Optimized expert helper** | **143.35** | **197.58** | --- ## What changed - **Complementary expert caches** Keep the primary and helper GPU expert caches complementary, reducing duplicated experts and CPU fallback. - **Lower inter-GPU transfer volume** Reduce weighted expert outputs on helper GPUs before returning results. - **Skip unnecessary CPU quantization** Avoid CPU activation quantization when all selected experts are served by GPUs. - **Automatic helper-cache sizing** Add `auto` sizing based on available VRAM, removing the need to manually tune expert counts. The existing primary/helper GPU layout remains unchanged.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。