Pull requests / #880

#880 layer split: the stage weight trim works under `auto` too 4X gain in the prefill in 4 way ring

closed · @gopinath87607 · 0 Kommentare · Auf GitHub

BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants

Beschreibung

`--trim-stage-weights` (PR #559) / `STRATA_STAGE_TRIM=1` (PR #639) gives a stage only its own layers' dense
weights, but it refuses to run unless `--layer-split` is written out: `auto` chooses the split points 590 lines
after the stages' weights are loaded, so by the time the split is known there is nothing left to trim.

Those weights now load after the search. The expert caches are sized later, from live free VRAM, so they collect
the freed bytes.

On the 4-way IQ3_S rig (2x RTX 3060 12 GB + 2x RTX 5060 8 GB), `--layer-split auto`, `--trim-stage-weights`:

    CUDA1 3650 -> 4608 expert slots, CUDA2 1241 -> 2485, CUDA3 533 -> 1772
    prompt chunk 768 -> 6144 tokens

CUDA0 still trims only with an explicit `--layer-split`. Its arena is read out to build the prompt path's PLE
tables, 590 lines before the split exists, so moving it is the larger half of this change. It is also the cheap
half to leave: measured, it costs 1%.

## The bug found in review, and the fix

Deferring the load also defers the allocations that come with it, and the search priced those from the pack
index instead of from `cudaMemGetInfo`. That price was wrong three ways:

1. **The arena was over-subtracted.** `stage_pool` was re-priced only `if (stage_trim)`, which is false under
   `auto` - so it held the whole pool while the stage really loaded a carved one.
2. **The native projections were not subtracted at all.** `NativeDense::load` allocates on the device, and
   nothing priced it. On this pack that is 2018.88 MiB over 300 matrices. This is the 1.92 GiB
   @paulhothersall measured on a P40 + RTX 3070 rig, and it is what made the planner tell the 8 GB card it had
   4.26 GiB free where v0.1.39 said 2.34 - so it picked **K=2 instead of K=16**, leaving the P40 on two layers.
3. **The `cudaMalloc` rounding was not priced.** `NativeDense::load` makes one `cudaMalloc` per matrix and the
   driver backs each call with 2 MiB granules, so the footprint is not the payload sum. Measured on the RTX 5060:
   300 buffers asking 2018.9 MiB cost **2400.0 MiB** - **381.1 MiB** of rounding the search could not see.

The fix, as @j-luwierski suggested: one per-stage footprint calculation, covering the arena and the native
projections, priced from the pack before the search and subtracted per candidate inside `predict` beside
`session_bytes`, memoised by layer range. `WeightTable::pool_bytes` reads `index.txt` and
`NativeDense::weight_bytes_for` reads the GGUF headers - neither allocates, so no device is needed and no stage
has to load early. `alloc_bytes` prices each matrix the way `cudaMalloc` backs it: a 2 MiB granule for every
matrix of a MiB or more, the payload below, because the driver packs the small ones into the slack of the rounded
ones. On this pack's 300 matrices that prices 2255 MiB against a real 2238, +0.8% - erring high on purpose, so the
card is left slightly less room than it has rather than slightly more.

## Measurements

Two devices (RTX 3060 12 GB + RTX 5060 8 GB), `--layer-split auto`:

| arm | split | the 8 GB card's priced footprint | really taken | prefill | decode |
|---|---|---|---|---|---|
| plain v0.1.39 (no deferral - ground truth) | **K=20** | 1.87 GiB | 3.63 GiB free | 107.2 | 34.8 |
| #880 before this commit | **K=6** | 4.06 GiB | 3.63 GiB free | *does not boot* | - |
| with this commit | **K=20** | 3721 MiB | 3706 MiB | 107.6 | 33.3 |

K=20 is what plain v0.1.39 picks with no deferral at all, so the fix reproduces the old planner rather than
inventing a new one - same split, same 435 expert-cache slots, same 256-token chunk, throughput within 4%.

With the trim on:

| arm | split | chunk | cache slots CUDA1 | prefill | decode | cache hit |
|---|---|---|---|---|---|---|
| #880 before this commit | **K=6** | 768 | 521 | 169.2 tok/s | 31.2 tok/s | 41.6% |
| with this commit | **K=17** | 2048 | 1014 | **364.9 tok/s** | **36.0 tok/s** | **61.5%** |

+116% prefill, +15% decode. The chunk follows the split: 768 -> 2048 tokens.

On 4 cards the split does not move - **K=10,19,34 in every arm** - and prefill is 884.0 tok/s against 878.8 for
the unfixed build, i.e. noise. That rig cannot show the flaw: with four stages no card's cap is the binding
constraint. It is a regression test, and it passes.

### Stacked on #887, which is where this actually bites

The paragraph above is why this PR's own four-card numbers look flat. #887 ("layer split: score every four-way
placement instead of guessing it") is the branch where it shows. #887 does **not** carry this bug - the deferral
comes from this PR - but it is what makes the four-way choice cap-bound, so a mis-priced card changes the answer
there instead of hiding inside a guess. 4 cards, `--layer-split auto --trim-stage-weights`:

| | split | cache slots CUDA1/2/3 | pairs held | predicted | prompt chunk |
|---|---|---|---|---|---|
| #887 + this PR, before this commit | K=10,28,39 | 4663 / 2696 / 1919 | 11966 (86.2% of the routed mass) | 85.8 ms | 6144 |
| #887 + this PR, with this commit | K=10,27,38 | 4787 / 2650 / 1891 | 12643 (88.3%) | 81.4 ms | 7168 |

**The two are worth reading as a pair, and testing as one.** #887 chooses between placements; this PR is what
prices the cards those placements are scored against.

My other two open PRs, #582 ("Serve tool result images") and #581 ("the Monitor's per-GPU cards"), are in the
build I run daily but are frontend and telemetry: they do not touch the split, the caches or the prompt chunk,
and they change none of these numbers.

## Self-check

After each stage loads, the code compares the bytes the search priced against the bytes that stage **really
took**, read from `cudaMemGetInfo` before and after the load. It warns on any under-pricing (the actual bug - the
card is handed more than it can hold) and on over-pricing past 5% (the granule rule runs ~0.8% high by design, so
real drift would stand out). It stayed silent in every arm above, including the two-device runs where the old
price was 2.49 GiB out.

## Known limits

- The 4-card rig cannot show the flaw. It needs one card much smaller than the others, which is what the
  two-device runs above are.
- The granule rule (2 MiB, ~1 MiB threshold) is measured on this NVIDIA driver and this one pack. HIP's allocator
  may round differently - @aswin-dot-R, if the split moves the wrong way on 2x gfx1201, say so and I will
  re-measure the constant.
- The P40 + RTX 3070 case itself has not been re-tested: I have no P40. @paulhothersall offered to re-run, and
  the log line to check is `layer split auto: CUDA1 ... GiB free before its weights, session and experts` - it
  read 2.34 GiB on v0.1.39, 4.26 under the old build, and should read 2.34 again.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Mehr auf der Site

Links zu Install, Modellen, Releases.