Issues / #509

#509 layer_split auto is 2.86x slower than a balanced K on prefill with a heterogeneous pair (3080 + V100, measured)

closed · @meelonjisoo-commits · 2 comments · View on GitHub

Multi-GPUAMD / HIPNVIDIA / CUDAModels & quants

Description

## Platform (all values measured on this machine)

| Item | Value |
|---|---|
| GPU 0 (main) | NVIDIA RTX 3080 10 GB — sm_86, gen4 x16, CPU-attached |
| GPU 1 | Tesla V100-PCIE-32GB — sm_70, **gen3 x4 (chipset)**, via riser |
| CPU | AMD Ryzen 5 5600X (6C/12T, AVX2, no AVX-512) |
| RAM | 92 GiB DDR4-2133 |
| OS / driver | Ubuntu, kernel 7.0.0-34, NVIDIA 580.178.04 |
| Strata | **0.1.35**, local dual-arch build (sm_86 + sm_70, 110 cubins), md5 `7872494f` |
| Model / config | Qwen3.8-Flash-Next IQ3_XXS (GSQ-RCO), ctx 163,840, `--kv q4_0`, `--expert-cache auto`, MTP draft |
| Probe | 32,368-token synthetic agent turn, cold (fresh process), `drop_caches` before every arm |
| Note | The V100 is below the documented 7.5 compute-capability floor, so this pair is off-spec — see the closing note for why it still may be interesting |

## What happens

`layer_split auto` picks a placement that is **2.86x slower on prompt processing** than an explicit balanced `K`:
**549.5 t/s vs 1572.7 t/s** on the same 32,368-token prompt, same engine, same build, same links. The single-card
baseline is 975.0 t/s — so the *default* split is **1.77x slower than not splitting**, while a balanced split is
**1.61x faster than not splitting**. Only the `K` differs.

| shape | placement | cold 32K prefill | turn 3 after an aux call | decode |
|---|---|---:|---:|---:|
| 3080 alone | 48 layers, one card | 975.0 t/s (33,199 ms) | 0.40 s | 37.4-55.5 |
| split, `layer_split auto` | **K=11** — V100 takes 37 of 48 layers | **549.5 t/s** (58,903 ms) | 58.5 s | 62.5-105.7 |
| split, one card, same K (`--layer-split 11 --split-device 0`) | hand-off control on a fast link | 972.4 t/s | 33.4 s | 39.4-64.8 |
| split, `--layer-split 24` | V100 takes 24 | 1269.2 t/s (25,504 ms) | 25.2 s | 57.2-95.9 |
| split, `--layer-split 30` | V100 takes 18 | **1572.7 t/s** (20,581 ms) | 20.3 s | 53.8-89.8 |

At K=30 the first card is the bottleneck (30 layers x ~0.69 s/layer = 20.7 s predicted, 20.6 s measured), so K~=29-30
is this pair's optimum and ~1600 t/s its ceiling. The V100's SM clock was sampled in every split arm
(median 1380 of 1380 MHz), so the numbers are not a clock artifact.

## Why auto lands on the slow placement

`MULTI_GPU.md` states auto "keeps the one whose caches would hold the most of the expert profile". That objective is
blind to the **per-layer cost of the card the layers land on**, and on a heterogeneous pair the cache-coverage
optimum and the time optimum diverge:

- auto gave the V100 37 of 48 layers (14,598 slots, ~99.2 % of the routed mass covered);
- but the V100 computes a layer ~1.34x slower than the 3080 (0.925 s vs 0.69 s per layer for a 32K pass, each
  measured alone on this box), so 37 layers there makes it the pipeline's bottleneck;
- balanced 30/18, both stages run concurrently and the prompt time is set by the slower stage's share: 1573 t/s.

The engine already estimates per-stage per-layer cost at startup for its own plan log
(`CUDA0 68 SMs at 1.78 GHz -> 0.60 ms per layer`, `CUDA1 80 SMs at 1.38 GHz -> 0.66 ms per layer`), so a
time-balanced auto would be `argmin_K max(K*t0, (48-K)*t1)` plus the hand-off instead of a coverage ranking. (Your own
`MULTI_GPU.md` table shows the same, smaller, gap: auto K=22 at 2,037 t/s against "best split (K=26)" at 2,039 /
2,357 t/s on a 5080+3090.)

## Two things it is NOT (both measured, so nobody re-chases them)

1. **Not the second card's link width.** A split stage's PCIe expert share is derived from its own probe
   (`bw>=20 -> 0.55 | bw<4 -> 0.00 | else 0.55*bw/26`); our V100 stage probes 2.8 GB/s -> 0.00. Forcing the policy a
   wider slot would select changed nothing at K=11: `--pcie-frac 0.14` -> 549.4, `0.28` -> 549.2, `0.55` -> 549.4 t/s.
   With a balanced K each stage's cache covers ~100 % of its own layers' profile, so almost nothing crosses the link
   during prefill at all.
2. **Not the two-stage mechanism.** `--layer-split 11 --split-device 0` (both stages on the one fast card, sharing
   weights/session/cache) costs **-0.3 %** against that card alone (972.4 vs 975.0 t/s).

## Reproducing

`--layer-split K` with `"gpu": [main, second]`, conversation-cache args removed (parking + split is rejected), a
32K-token cold prompt, `drop_caches` before each run, and the V100's SM clock sampled during the arm. The engine is
otherwise pristine 0.1.35: a dual-arch build is needed for the sm_70 card, which we did by lowering the CMake
compute-capability floor to 70 (upstream's below-7.5 community build option, #236, is the same idea). Any pair whose
two cards differ in per-layer cost (a 24 GB + 8 GB mix, a card in a narrow slot, an older card) should show the same
divergence; happy to re-measure a suggested placement on this machine.

## Related

- **#448** (closed, fixed in 0.1.35): the small-second-card 512-token chunk cap — a *different* cause of slow split
  prompts, and it was our first suspect for this. It is not this: our slow arm (K=11) and our fast arm (K=30) run the
  same 0.1.35 and the same 512-cap-free code path, and the only difference is `K`.
- **#305 / #236**: sm_70 / below-7.5 support path — this pair (RTX 3080 + V100) is an off-spec, heterogeneous
  combination, which is exactly where auto's coverage objective and real time diverge most.
- **#490** (open) is a different heterogeneous idea (the MTP draft head on the second card) and does not overlap
  with this placement question.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.