Pull requests / #584
#584 Split placement search
closed · @gopinath87607 · 0 Kommentare · Auf GitHub
Server & APIMulti-GPUAMD / HIPNVIDIA / CUDA
Beschreibung
## What this is **Depends on the weight carve (PR #580, `weight-carve-0138`), and is one commit on top of it.** Upstream only takes a PR whose base branch is in `Niko1221/Strata`, so this opens against `main` and the compare shows A's five commits too — **only the last one, "score every four-way placement instead of guessing it", is this PR**. Once A is merged I'll rebase this onto `main`, and then it is a single commit. A four-GPU run picked its split from a formula that reads only `layer_ms`. `capr` — the free VRAM that decides how many experts a stage can actually hold — never entered it, so the search could hand the most layers to the card with the least room and never notice. All C(L−3,3) placements are now scored: 15,180 at 48 layers. Measured on this 3060/3060/5060/5060 box **before the weight carve landed**, the heuristic's answer: | | | |---|---| | CUDA1 | 1,841 MiB no cache could use | | CUDA2 | 15 layers it could hold a third of | | across the two 3060s | ~20 GiB stranded | **After the carve that gap is much smaller**, and this PR should be sized off the second table, not the first: a stage no longer holds the 3,485.66 MiB of dense weights it will never run, so there is far less VRAM to redeploy — the stranding the 20 GiB came from is mostly gone. Same box, same flags (`--layer-split auto --expert-cache auto --prefill auto:32768`), the carve in both arms, so the split is the only variable: | | heuristic | the search | |---|---|---| | split | 0-13 / 14-28 / 29-37 / 38-47 | 0-9 / 10-20 / 21-34 / 35-47 | | CUDA0 cache | 2,974 slots | **3,049** | | CUDA1 cache | 2,552 | **3,019** | | CUDA2 cache | 4,608 | **4,524** | | CUDA3 cache | 3,927 | **3,888** | | expert slots | 14,061 | **14,480** | | expert cache | 27,727 MiB | **28,242 MiB** | **+419 slots, +3.0%** — a third of the pre-carve figure, and the honest number to merge on. What is still wrong is visible in the per-stage counts: CUDA1 held 2,552 of the 7,680 profiled pairs its range could see, while CUDA2 was saturated at 4,608 of 4,608. The search moves layers off the card that cannot hold them, and CUDA1 ends up covering 3,019 of 5,632. ## What makes it affordable A stage's fill depends only on its own layer range and its own free VRAM — each ranked pair goes to the one stage that owns its layer, and each stage stops at its first pair that does not fit. A placement's held mass is therefore a sum over four independently memoised range walks: **1,121 walks instead of 15,180**, measured at **14 ms** against 740 ms to evaluate the same answer directly. ## The mass curve had to move with it Scoring the placements is not enough on its own, because the cost model's mass curve was wrong here. `mass[r]` was `(r+1)^-1.2`, fitted to the 5080+3090 sweep. It is still right on that rig (99.35% predicted at 20,000 held, 99.4% measured) and badly wrong on this one: it claims the top 5,805 pairs hold 94.9% of the routed mass where the rig measures **69.2%** over four boots and 14.9M decode lookups. No exponent reconciles the two, because the curve is a property of the model's routing, and the profile file carries only a ranking — no counts — so it cannot be read off the data. What both rigs do agree on is that the hit rate follows the held **count** almost alone: 14,541 held measured 92.2%, 14,384 held measured 89.8%. The mass is therefore the increment of `H(f) = 1 − (1 − f)^b` over the held fraction, whose prefixes sum to exactly `H(N/total)`. `b = 3` is chosen to leave the validated regime where it was: at f = 0.8138 it gives 99.35%, the same as the old fit, so the **2-GPU placements are untouched**, while at this rig's f = 0.59 it gives 93.2% against 92.2% measured. `STRATA_SPLIT_COVER_B` overrides `b` for a model whose routing is sharper. Because this changes the **chosen split**, it is its own PR: it would break the weight-carve PR's "identical to before" claim if it rode along with it. ## Also `tools/make_profile.py` gains a note: `take()` skips a pair it has already ranked, and the shipped base ranks all 24,576, so with the default `--base` no trace can move a pair — the run prints "0 from the traces" and the output is the base again. The ordering in use is the shipped one, not this model's, which is why the cost model needed the measured curve above rather than another fit. ## Testing The placement search is exercised by the existing split-search paths; `serve`'s Python tests are unaffected. On the rig, the four-way split it chooses is compared against the heuristic's on the same box (`INFO`'s slot counts, and the per-stage `layer split:` lines) — the two tables above, one run of each branch's binary with the flags in their captions. It is a run of the real engine on the real pack, not a unit test: there is no multi-GPU integration test in this repo. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Mehr auf der Site
Links zu Install, Modellen, Releases.