Pull requests / #584

#584 Split placement search

closed · @gopinath87607 · 0 comentários · No GitHub

Server & APIMulti-GPUAMD / HIPNVIDIA / CUDA

Descrição

## What this is

**Depends on the weight carve (PR #580, `weight-carve-0138`), and is one commit on top of it.** Upstream
only takes a PR whose base branch is in `Niko1221/Strata`, so this opens against `main` and the compare
shows A's five commits too — **only the last one, "score every four-way placement instead of guessing
it", is this PR**. Once A is merged I'll rebase this onto `main`, and then it is a single commit.

A four-GPU run picked its split from a formula that reads only `layer_ms`. `capr` — the free VRAM
that decides how many experts a stage can actually hold — never entered it, so the search could
hand the most layers to the card with the least room and never notice. All C(L−3,3) placements are
now scored: 15,180 at 48 layers.

Measured on this 3060/3060/5060/5060 box **before the weight carve landed**, the heuristic's answer:

| | |
|---|---|
| CUDA1 | 1,841 MiB no cache could use |
| CUDA2 | 15 layers it could hold a third of |
| across the two 3060s | ~20 GiB stranded |

**After the carve that gap is much smaller**, and this PR should be sized off the second table, not
the first: a stage no longer holds the 3,485.66 MiB of dense weights it will never run, so there is
far less VRAM to redeploy — the stranding the 20 GiB came from is mostly gone. Same box, same flags
(`--layer-split auto --expert-cache auto --prefill auto:32768`), the carve in both arms, so the split
is the only variable:

| | heuristic | the search |
|---|---|---|
| split | 0-13 / 14-28 / 29-37 / 38-47 | 0-9 / 10-20 / 21-34 / 35-47 |
| CUDA0 cache | 2,974 slots | **3,049** |
| CUDA1 cache | 2,552 | **3,019** |
| CUDA2 cache | 4,608 | **4,524** |
| CUDA3 cache | 3,927 | **3,888** |
| expert slots | 14,061 | **14,480** |
| expert cache | 27,727 MiB | **28,242 MiB** |

**+419 slots, +3.0%** — a third of the pre-carve figure, and the honest number to merge on. What is
still wrong is visible in the per-stage counts: CUDA1 held 2,552 of the 7,680 profiled pairs its
range could see, while CUDA2 was saturated at 4,608 of 4,608. The search moves layers off the card
that cannot hold them, and CUDA1 ends up covering 3,019 of 5,632.

## What makes it affordable

A stage's fill depends only on its own layer range and its own free VRAM — each ranked pair goes to
the one stage that owns its layer, and each stage stops at its first pair that does not fit. A
placement's held mass is therefore a sum over four independently memoised range walks: **1,121
walks instead of 15,180**, measured at **14 ms** against 740 ms to evaluate the same answer
directly.

## The mass curve had to move with it

Scoring the placements is not enough on its own, because the cost model's mass curve was wrong
here. `mass[r]` was `(r+1)^-1.2`, fitted to the 5080+3090 sweep. It is still right on that rig
(99.35% predicted at 20,000 held, 99.4% measured) and badly wrong on this one: it claims the top
5,805 pairs hold 94.9% of the routed mass where the rig measures **69.2%** over four boots and
14.9M decode lookups.

No exponent reconciles the two, because the curve is a property of the model's routing, and the
profile file carries only a ranking — no counts — so it cannot be read off the data. What both rigs
do agree on is that the hit rate follows the held **count** almost alone: 14,541 held measured
92.2%, 14,384 held measured 89.8%.

The mass is therefore the increment of `H(f) = 1 − (1 − f)^b` over the held fraction, whose
prefixes sum to exactly `H(N/total)`. `b = 3` is chosen to leave the validated regime where it was:
at f = 0.8138 it gives 99.35%, the same as the old fit, so the **2-GPU placements are untouched**,
while at this rig's f = 0.59 it gives 93.2% against 92.2% measured. `STRATA_SPLIT_COVER_B`
overrides `b` for a model whose routing is sharper.

Because this changes the **chosen split**, it is its own PR: it would break the weight-carve PR's
"identical to before" claim if it rode along with it.

## Also

`tools/make_profile.py` gains a note: `take()` skips a pair it has already ranked, and the shipped
base ranks all 24,576, so with the default `--base` no trace can move a pair — the run prints
"0 from the traces" and the output is the base again. The ordering in use is the shipped one, not
this model's, which is why the cost model needed the measured curve above rather than another fit.

## Testing

The placement search is exercised by the existing split-search paths; `serve`'s Python tests are
unaffected. On the rig, the four-way split it chooses is compared against the heuristic's on the
same box (`INFO`'s slot counts, and the per-stage `layer split:` lines) — the two tables above, one
run of each branch's binary with the flags in their captions. It is a run of the real engine on
the real pack, not a unit test: there is no multi-GPU integration test in this repo.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

No site

Links install, modelos, releases.