Pull requests / #887
#887 layer split: score every four-way placement instead of guessing it
closed · @gopinath87607 · 0 commentaires · Sur GitHub
BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants
Description
**What this does.** With four GPUs, `--layer-split auto` picked its split from a formula that reads only `layer_ms`
(SMs x clock) and shares the layers in proportion to speed. The free VRAM left for expert caches - the thing that
decides how much of the model's routing a stage can hold - never entered it, so the search could hand the most layers
to the card with the least room and never notice. Every placement is now scored.
**All C(L-3,3) placements are scored** - 15,180 at 48 layers. What makes that affordable: a stage's fill depends
only on its own layer range and its own free VRAM. `predict` already hands each ranked pair to the one stage that owns
its layer and stops each stage at its first pair that does not fit, so the four fills are independent and a
placement's held mass is a SUM over its four ranges, so a range is walked once and memoised: at 48 layers the
enumeration scores 15,180 placements with 2,068 walks (counted from the loop bounds; timed on the rig below).
**The cost model's mass curve had to change with it.** `mass[r]` was `(r+1)^-1.2`, a fit to the 5080+3090 sweep. That
fit is still right on the rig it was fitted to - at 20,000 pairs held it predicts 99.35% of the routed mass and that
rig measured 99.4% - and wrong on IQ3_S: it claims the top 5,805 pairs hold 94.9% of the mass where a 4-way rig
measures 69.2% over four boots and 14.9M decode lookups. No exponent reconciles the two, because the curve is a
property of the model's routing, and the profile file carries only a ranking, with no counts, to read it off. What
both rigs agree on is that the hit rate follows the held COUNT and almost nothing else, so the mass is now the
increment of `H(f) = 1 - (1 - f)^b` over the held fraction, whose prefixes sum to exactly `H(N/total)`. `b = 3` is
chosen to leave the validated regime where it was - at `f = 0.8138` it gives 99.35%, the same as the old fit - while
at `f = 0.59` it gives 93.2% against 92.2% measured. Below `f ~ 0.5` the two curves diverge, and that is the point:
it is exactly where the old one told the search that stranding a card's VRAM costs almost nothing.
`STRATA_SPLIT_COVER_B` overrides `b`.
The curve applies to the two- and three-GPU searches as well, not only the new four-way one - `b = 3` is what keeps
them where they were at the held fractions the old fit had been validated on.
**The startup line now says how many placements were scored**, and on the 4-way IQ3_S rig it prints this:
layer split auto: K=14,31,41 - predicted 114.6 ms per decode window; the caches hold 8620 of 24576
profiled pairs (~71.9% of the routed mass), best of 15180 placements
**This supersedes #584**, which is the same search written on top of a weight carve. v0.1.39 now ships its own carve
(`--trim-stage-weights`, #559, and `STRATA_STAGE_TRIM=1`, #639) built on a different mechanism - `add_foreign` fills
the loader's skip set instead of adding a ranged `WeightTable` API - so #584's five carve commits cannot be merged
beside it: they would be a second arena-sizing path for the same bytes. This carries the search alone.
**One thing the port does differently, deliberately.** On the branch the carve reached `--layer-split auto`, so the
search had to take each stage's own dense-weight bytes off its room itself. On v0.1.39 the search only runs under
`auto`, where `--trim-stage-weights` is refused and every card still holds the whole model's dense weights - so
`cap[i]` is already net of them, and the room a stage has is what `predict` already subtracted: its session. The
enumeration prices exactly that. When the carve does reach `auto`, this is the line to revisit.
**Measured on the rig.** 4-way IQ3_S (2x RTX 3060 12 GB + 2x RTX 5060 8 GB), `--layer-split auto`,
`--expert-cache auto`, `--prefill auto:32768`, no `--trim-stage-weights` - the same config file, the same flags, the
only difference is the binary. Two boots per arm, back to back; prefill is one cold 10,043-token prompt asked for 1
token, decode a 104-token prompt asked for 300.
| arm | split | chunk / ring | caches CUDA1/2/3 | prefill | decode |
|---|---|---|---|---|---|
| v0.1.39 as shipped | K=10,19,34 | 768 / 8 slots | 3650 / 1241 / 533 | 425.4, 425.3 tok/s | 43.8, 45.1 tok/s |
| this branch | K=14,31,41 | 1280 / 26 slots | 3316 / 1272 / 572 | 595.3, 598.3 tok/s | 51.1, 51.8 tok/s |
**+40% prefill** (two runs: +39.9%, +40.7%) and **+16% decode** (+16.7%, +14.9%). The split is deterministic - both
boots of an arm print the same K.
**Why it wins.** The shipped split hands 15 layers to CUDA2 and 14 to CUDA3 - the two 8 GB cards - and leaves them
1,241 and 533 expert-cache slots. The search moves the layers to the 12 GB cards instead (CUDA2 10 layers, CUDA3 7),
so even though it ends up with *fewer* total slots (5,160 against 5,424) the coverage per layer is better: decode
expert-cache hit 83.3% against 74.7%, and the starved cards now lend enough for a 1,280-token prompt chunk where
they capped it at 768.
**A number that looks worse and is not.** The printed "predicted 114.6 ms per decode window" is well above the
shipped line's 63.9 ms, and the two are not comparable: they come from different mass curves. The old curve claims
this split holds 95.5% of the routed mass; the rig measured 74.7% for it. The new one claims 71.9% where the rig
measured 83.3%. Both are estimates and both are wrong; the difference is which way. The measured tok/s is the
comparison to read, not the predicted ms.
**What this measurement does not separate.** The mass curve and the search are one commit, and both arms here
differ from the base by that commit alone - so the +40%/+16% is the pair, not the search on its own. A third arm
(the new curve with the old proportional four-way guess) would split them; it needs a temporary switch and another
build, and is worth doing before this merges.
🤖 Generated with [Claude Code](https://claude.com/claude-code)Sur le site
Liens install, modèles, releases.