Issues / #509
#509 layer_split auto is 2.86x slower than a balanced K on prefill with a heterogeneous pair (3080 + V100, measured)
closed · @meelonjisoo-commits · 2 comentários · No GitHub
Multi-GPUAMD / HIPNVIDIA / CUDAModels & quants
Descrição
## Platform (all values measured on this machine) | Item | Value | |---|---| | GPU 0 (main) | NVIDIA RTX 3080 10 GB — sm_86, gen4 x16, CPU-attached | | GPU 1 | Tesla V100-PCIE-32GB — sm_70, **gen3 x4 (chipset)**, via riser | | CPU | AMD Ryzen 5 5600X (6C/12T, AVX2, no AVX-512) | | RAM | 92 GiB DDR4-2133 | | OS / driver | Ubuntu, kernel 7.0.0-34, NVIDIA 580.178.04 | | Strata | **0.1.35**, local dual-arch build (sm_86 + sm_70, 110 cubins), md5 `7872494f` | | Model / config | Qwen3.8-Flash-Next IQ3_XXS (GSQ-RCO), ctx 163,840, `--kv q4_0`, `--expert-cache auto`, MTP draft | | Probe | 32,368-token synthetic agent turn, cold (fresh process), `drop_caches` before every arm | | Note | The V100 is below the documented 7.5 compute-capability floor, so this pair is off-spec — see the closing note for why it still may be interesting | ## What happens `layer_split auto` picks a placement that is **2.86x slower on prompt processing** than an explicit balanced `K`: **549.5 t/s vs 1572.7 t/s** on the same 32,368-token prompt, same engine, same build, same links. The single-card baseline is 975.0 t/s — so the *default* split is **1.77x slower than not splitting**, while a balanced split is **1.61x faster than not splitting**. Only the `K` differs. | shape | placement | cold 32K prefill | turn 3 after an aux call | decode | |---|---|---:|---:|---:| | 3080 alone | 48 layers, one card | 975.0 t/s (33,199 ms) | 0.40 s | 37.4-55.5 | | split, `layer_split auto` | **K=11** — V100 takes 37 of 48 layers | **549.5 t/s** (58,903 ms) | 58.5 s | 62.5-105.7 | | split, one card, same K (`--layer-split 11 --split-device 0`) | hand-off control on a fast link | 972.4 t/s | 33.4 s | 39.4-64.8 | | split, `--layer-split 24` | V100 takes 24 | 1269.2 t/s (25,504 ms) | 25.2 s | 57.2-95.9 | | split, `--layer-split 30` | V100 takes 18 | **1572.7 t/s** (20,581 ms) | 20.3 s | 53.8-89.8 | At K=30 the first card is the bottleneck (30 layers x ~0.69 s/layer = 20.7 s predicted, 20.6 s measured), so K~=29-30 is this pair's optimum and ~1600 t/s its ceiling. The V100's SM clock was sampled in every split arm (median 1380 of 1380 MHz), so the numbers are not a clock artifact. ## Why auto lands on the slow placement `MULTI_GPU.md` states auto "keeps the one whose caches would hold the most of the expert profile". That objective is blind to the **per-layer cost of the card the layers land on**, and on a heterogeneous pair the cache-coverage optimum and the time optimum diverge: - auto gave the V100 37 of 48 layers (14,598 slots, ~99.2 % of the routed mass covered); - but the V100 computes a layer ~1.34x slower than the 3080 (0.925 s vs 0.69 s per layer for a 32K pass, each measured alone on this box), so 37 layers there makes it the pipeline's bottleneck; - balanced 30/18, both stages run concurrently and the prompt time is set by the slower stage's share: 1573 t/s. The engine already estimates per-stage per-layer cost at startup for its own plan log (`CUDA0 68 SMs at 1.78 GHz -> 0.60 ms per layer`, `CUDA1 80 SMs at 1.38 GHz -> 0.66 ms per layer`), so a time-balanced auto would be `argmin_K max(K*t0, (48-K)*t1)` plus the hand-off instead of a coverage ranking. (Your own `MULTI_GPU.md` table shows the same, smaller, gap: auto K=22 at 2,037 t/s against "best split (K=26)" at 2,039 / 2,357 t/s on a 5080+3090.) ## Two things it is NOT (both measured, so nobody re-chases them) 1. **Not the second card's link width.** A split stage's PCIe expert share is derived from its own probe (`bw>=20 -> 0.55 | bw<4 -> 0.00 | else 0.55*bw/26`); our V100 stage probes 2.8 GB/s -> 0.00. Forcing the policy a wider slot would select changed nothing at K=11: `--pcie-frac 0.14` -> 549.4, `0.28` -> 549.2, `0.55` -> 549.4 t/s. With a balanced K each stage's cache covers ~100 % of its own layers' profile, so almost nothing crosses the link during prefill at all. 2. **Not the two-stage mechanism.** `--layer-split 11 --split-device 0` (both stages on the one fast card, sharing weights/session/cache) costs **-0.3 %** against that card alone (972.4 vs 975.0 t/s). ## Reproducing `--layer-split K` with `"gpu": [main, second]`, conversation-cache args removed (parking + split is rejected), a 32K-token cold prompt, `drop_caches` before each run, and the V100's SM clock sampled during the arm. The engine is otherwise pristine 0.1.35: a dual-arch build is needed for the sm_70 card, which we did by lowering the CMake compute-capability floor to 70 (upstream's below-7.5 community build option, #236, is the same idea). Any pair whose two cards differ in per-layer cost (a 24 GB + 8 GB mix, a card in a narrow slot, an older card) should show the same divergence; happy to re-measure a suggested placement on this machine. ## Related - **#448** (closed, fixed in 0.1.35): the small-second-card 512-token chunk cap — a *different* cause of slow split prompts, and it was our first suspect for this. It is not this: our slow arm (K=11) and our fast arm (K=30) run the same 0.1.35 and the same 512-cap-free code path, and the only difference is `K`. - **#305 / #236**: sm_70 / below-7.5 support path — this pair (RTX 3080 + V100) is an off-spec, heterogeneous combination, which is exactly where auto's coverage objective and real time diverge most. - **#490** (open) is a different heterogeneous idea (the MTP draft head on the second card) and does not overlap with this placement question.
No site
Links install, modelos, releases.