Pull requests / #794
#794 split: the auto placement sees each stage's PCIe link
closed · @chengshenyangtai · 0 commentaires · Sur GitHub
BenchmarksSetup & installMulti-GPUNVIDIA / CUDA
Description
## What this fixes
`--layer-split auto` picks the split by a predicted window time: layer compute plus a miss cost for experts no cache
holds. That miss cost (`190 ms per unit of routed mass`) is **one constant for every stage** — it was fitted on a pair
of equal x16 cards, so the search cannot see an asymmetric pair and spreads layers by compute and cache size alone.
Measured case: 2x RTX 4090 D, the second card on a **physical x1 link** (probed: 23.9 GB/s on the x16 card vs 1.6 GB/s
on the x1 card). The prompt path streams every missed expert over *its stage's* link, so the layers given to the x1 card
dominate every chunk:
| | placement | 29K-token prompt, fresh reads |
|---|---|---|
| search as-is | K=22 (26 layers on the x1 card) | ~350 tok/s |
| with this change | K=38 (the x1 card's cache holds all of its layers' pairs — zero streaming there) | ~1410 tok/s |
Same machine, same model (Qwen3.8-Flash-Next UD-Q4_K_XL), 3 fresh real-text prompts per arm, engine-reported timings.
The old placement's per-chunk streaming on x1 is ~26 GB ÷ 1.6 GB/s ≈ 17 s; the observed chunk time matched.
## How
- Probe each stage's link inside the search (`probe_pcie_h2d_gbps`). The stage-setup probes are skipped when
`--pcie-frac` is given, and the search needs the reading either way.
- Scale each stage's miss term by `clamp(20 GB/s / bw, 1, 64)`. 20 GB/s is `pcie_frac_for_gbps`' x16 reference; the
lower clamp keeps a symmetric machine on the exact formula the sweep's K was fitted with (all scales 1.0 → the objective
is unchanged).
- The objective's miss term becomes per-stage: `Σ_i miss_ms × scale_i × (missed mass of stage i)`, with the per-stage
held mass now accumulated in the placement fill.
Log lines on that machine:
layer split auto: CUDA0 link 23.9 GB/s -> miss cost x1.00
layer split auto: CUDA1 link 1.6 GB/s -> miss cost x12.15
layer split auto: K=38 - predicted 17.1 ms per decode window; the caches hold 10451 of 24576 profiled pairs (~97.4% of
the routed mass)
## Limits / notes
- The 190 ms constant mixes the decode CPU-pool cost with the prompt link cost; scaling the whole per-stage term is the
first-order fix. Splitting the constant into its two parts would need a fresh fit on both a narrow and a wide link —
happy to do that if you want it.
- Symmetric pairs are untouched: an x16 gen4/5 link reads 23-28 GB/s, so its scale is exactly 1.0.
- `--pcie-frac` still overrides the decode share exactly as before; this only changes the *placement search*'s cost
model.Sur le site
Liens install, modèles, releases.