Pull requests / #223

#223 Multi gpu: Layer split: pinned arena, lent prompt buffers, tested planner (2 x 2080 Ti decode +34%, 29K prompts +61%)

closed · @giostrives · 0 评论 · 在 GitHub 查看

BenchmarksSetup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

描述

On 2 x RTX 2080 Ti the layer split was barely faster than one card for decode and about 30% slower for prompts. This branch fixes how the split is set up. On this setup the split is now the faster configuration for both decode and long prompts.

Changes

1. layer split: pin the whole expert arena outside Windows (bb321e5). The 8 GiB pin cap is a WDDM workaround and now applies on Windows only. With the cap on Linux, only the first ~10 layers' misses could take the PCIe share, and the prompt path staged ~80% of the streamed experts through host copies.
2. layer split: every card lends its prompt buffers from its own cache (b578d5e). Each card used to reserve 2048-token buffers for the whole session (~1.5 GiB a card), and prompts ran in 2048-token chunks. Now every card lends its buffers from its expert cache, and all cards read the same chunk, the largest every card can lend.
3. layer split auto: the planner in layer_split.hpp, tested (251fb73). This is layer_split::Planner, with unit tests in src/program/layer_split_test.cpp. It adds a hand-off term per later card and makes the routing-curve exponent a parameter. The commit also fixes the later card's reserve, which counted the drafter twice: it was 1 GiB and is now 128 MiB.
4. bench: the layer split on 2 x RTX 2080 Ti (6bc110a). Harness, raw results and tables in bench/results/2026-09-30-layer-split-2080ti/, plus a section in docs/MULTI_GPU.md. Most of the added lines are here.
 

Related

The first commit is the same fix as #209 (cudaFuncGetName guarded for CUDA < 12.3). If #209 lands first, I'll drop the commit and rebase. Otherwise this MR carries it, and #209 can be closed as superseded.

Results

Qwen3.8-Flash-Next IQ2_XS, 2 x RTX 2080 Ti 11 GB, i7-7700K, 32K context, int8 KV, --spec 4. Same harness for every row, and every needle was found.

| Setup | Decode tok/s | Hits | Prompt 2.7K tok/s | Prompt 8K tok/s | Prompt 16K tok/s | Prompt 29K tok/s |
|---|---:|---:|---:|---:|---:|---:|
| 1 GPU | 33.9 | 64% | 618 | 712 | 795 | 801 |
| 2 GPUs, before (K=31) | 35.8 | 69% | 391 | 497 | 537 | 535 |
| 2 GPUs, after (auto K=25) | 45.3-47.3 | 80-82% | 609-627 | 826-829 | 1,120-1,126 | 1,290 |


- Decode on this machine is bound by the CPU pool, so the split pays through fewer CPU misses (hit rate 64% → 80-82%).
- The K sweep agrees with auto's pick: K=18 41.4, K=25 45-47, K=31 45.2, K=36 42.2 tok/s.
- On 0.1.27: one card 35.7 tok/s decode, the split 47.3. At 29K, prompts read 832 vs 1,304 tok/s.

Testing

- layer_split_test covers the planner. In the published tree it builds with g++ -std=c++20 -Iinclude src/program/layer_split_test.cpp && ./a.out.
- The full -DSTRATA_BUILD_TESTS=ON build does not work in the published tree, because some sources are not shipped. This is not new to this branch.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。