Issues / #509

#509 layer_split auto is 2.86x slower than a balanced K on prefill with a heterogeneous pair (3080 + V100, measured)

closed · @meelonjisoo-commits · 2 comentarios · En GitHub

Multi-GPUAMD / HIPNVIDIA / CUDAModels & quants

Descripción

## Platform (all values measured on this machine)

| Item | Value |
|---|---|
| GPU 0 (main) | NVIDIA RTX 3080 10 GB — sm_86, gen4 x16, CPU-attached |
| GPU 1 | Tesla V100-PCIE-32GB — sm_70, **gen3 x4 (chipset)**, via riser |
| CPU | AMD Ryzen 5 5600X (6C/12T, AVX2, no AVX-512) |
| RAM | 92 GiB DDR4-2133 |
| OS / driver | Ubuntu, kernel 7.0.0-34, NVIDIA 580.178.04 |
| Strata | **0.1.35**, local dual-arch build (sm_86 + sm_70, 110 cubins), md5 `7872494f` |
| Model / config | Qwen3.8-Flash-Next IQ3_XXS (GSQ-RCO), ctx 163,840, `--kv q4_0`, `--expert-cache auto`, MTP draft |
| Probe | 32,368-token synthetic agent turn, cold (fresh process), `drop_caches` before every arm |
| Note | The V100 is below the documented 7.5 compute-capability floor, so this pair is off-spec — see the closing note for why it still may be interesting |

## What happens

`layer_split auto` picks a placement that is **2.86x slower on prompt processing** than an explicit balanced `K`:
**549.5 t/s vs 1572.7 t/s** on the same 32,368-token prompt, same engine, same build, same links. The single-card
baseline is 975.0 t/s — so the *default* split is **1.77x slower than not splitting**, while a balanced split is
**1.61x faster than not splitting**. Only the `K` differs.

| shape | placement | cold 32K prefill | turn 3 after an aux call | decode |
|---|---|---:|---:|---:|
| 3080 alone | 48 layers, one card | 975.0 t/s (33,199 ms) | 0.40 s | 37.4-55.5 |
| split, `layer_split auto` | **K=11** — V100 takes 37 of 48 layers | **549.5 t/s** (58,903 ms) | 58.5 s | 62.5-105.7 |
| split, one card, same K (`--layer-split 11 --split-device 0`) | hand-off control on a fast link | 972.4 t/s | 33.4 s | 39.4-64.8 |
| split, `--layer-split 24` | V100 takes 24 | 1269.2 t/s (25,504 ms) | 25.2 s | 57.2-95.9 |
| split, `--layer-split 30` | V100 takes 18 | **1572.7 t/s** (20,581 ms) | 20.3 s | 53.8-89.8 |

At K=30 the first card is the bottleneck (30 layers x ~0.69 s/layer = 20.7 s predicted, 20.6 s measured), so K~=29-30
is this pair's optimum and ~1600 t/s its ceiling. The V100's SM clock was sampled in every split arm
(median 1380 of 1380 MHz), so the numbers are not a clock artifact.

## Why auto lands on the slow placement

`MULTI_GPU.md` states auto "keeps the one whose caches would hold the most of the expert profile". That objective is
blind to the **per-layer cost of the card the layers land on**, and on a heterogeneous pair the cache-coverage
optimum and the time optimum diverge:

- auto gave the V100 37 of 48 layers (14,598 slots, ~99.2 % of the routed mass covered);
- but the V100 computes a layer ~1.34x slower than the 3080 (0.925 s vs 0.69 s per layer for a 32K pass, each
  measured alone on this box), so 37 layers there makes it the pipeline's bottleneck;
- balanced 30/18, both stages run concurrently and the prompt time is set by the slower stage's share: 1573 t/s.

The engine already estimates per-stage per-layer cost at startup for its own plan log
(`CUDA0 68 SMs at 1.78 GHz -> 0.60 ms per layer`, `CUDA1 80 SMs at 1.38 GHz -> 0.66 ms per layer`), so a
time-balanced auto would be `argmin_K max(K*t0, (48-K)*t1)` plus the hand-off instead of a coverage ranking. (Your own
`MULTI_GPU.md` table shows the same, smaller, gap: auto K=22 at 2,037 t/s against "best split (K=26)" at 2,039 /
2,357 t/s on a 5080+3090.)

## Two things it is NOT (both measured, so nobody re-chases them)

1. **Not the second card's link width.** A split stage's PCIe expert share is derived from its own probe
   (`bw>=20 -> 0.55 | bw<4 -> 0.00 | else 0.55*bw/26`); our V100 stage probes 2.8 GB/s -> 0.00. Forcing the policy a
   wider slot would select changed nothing at K=11: `--pcie-frac 0.14` -> 549.4, `0.28` -> 549.2, `0.55` -> 549.4 t/s.
   With a balanced K each stage's cache covers ~100 % of its own layers' profile, so almost nothing crosses the link
   during prefill at all.
2. **Not the two-stage mechanism.** `--layer-split 11 --split-device 0` (both stages on the one fast card, sharing
   weights/session/cache) costs **-0.3 %** against that card alone (972.4 vs 975.0 t/s).

## Reproducing

`--layer-split K` with `"gpu": [main, second]`, conversation-cache args removed (parking + split is rejected), a
32K-token cold prompt, `drop_caches` before each run, and the V100's SM clock sampled during the arm. The engine is
otherwise pristine 0.1.35: a dual-arch build is needed for the sm_70 card, which we did by lowering the CMake
compute-capability floor to 70 (upstream's below-7.5 community build option, #236, is the same idea). Any pair whose
two cards differ in per-layer cost (a 24 GB + 8 GB mix, a card in a narrow slot, an older card) should show the same
divergence; happy to re-measure a suggested placement on this machine.

## Related

- **#448** (closed, fixed in 0.1.35): the small-second-card 512-token chunk cap — a *different* cause of slow split
  prompts, and it was our first suspect for this. It is not this: our slow arm (K=11) and our fast arm (K=30) run the
  same 0.1.35 and the same 512-cap-free code path, and the only difference is `K`.
- **#305 / #236**: sm_70 / below-7.5 support path — this pair (RTX 3080 + V100) is an off-spec, heterogeneous
  combination, which is exactly where auto's coverage objective and real time diverge most.
- **#490** (open) is a different heterogeneous idea (the MTP draft head on the second card) and does not overlap
  with this placement question.

En el sitio

Enlaces a install, modelos, releases.