Pull requests / #390

#390 Multi-GPU: a --layer-split stage holds only its own layers' weights, and every card's leftover VRAM spills into a helper expert tier (4-way IQ3_S: 9,269 -> 14,172 expert slots, prefill chunk 2,048 -> 4,096)

closed · @gopinath87607 · 0 comentarios · En GitHub

Server & APIMulti-GPUNVIDIA / CUDAModels & quants

Descripción

## What this is

Three things a `--layer-split` run needed, plus the Monitor work that made the first one visible.

The headline is the **carve**: every stage used to load the whole model's dense weights. Then the
VRAM those weights were wasting is given back — first to each stage's own expert cache, and then,
for whatever is still left over, to a **helper tier** that any stage's misses can hit.

## 1. A stage holds only its own layers' weights

Each `--layer-split` stage loaded the **whole model's** dense weights: on the 4-GPU IQ3_S rig,
3,485.66 MiB per card — 1,466.78 pack arena + 2,018.88 native dense — about 10.2 GiB across four
cards, holding layers that stage will never run. A stage running 9 of 48 layers still held all 48.

Weights now get the treatment the session already got. `session_bytes(g, cells, k, lo, hi)` is pure
arithmetic, so the split search *prices* a candidate layer range before allocating it; weights were
the one per-stage resource that never got that. The carve extends the same rule: a layer-range
predicate on the loader's existing skip/compact path, and the per-stage load moved down to after
`st.lb/le` are set. One rule produces both the priced size and the loaded layout, so they cannot
disagree.

Measured, 4-way IQ3_S (split 10,19,34 auto):

| | before | after |
|---|---|---|
| expert slots | 9,269 | **14,172** |
| expert cache | 17,388 MiB | **26,892 MiB** |
| CUDA0 slots | 3,657 | **5,120** |
| auto prefill chunk | 2,048 | **4,096** |

Per-card reclaim: CUDA0 2.56 GiB, CUDA1 2.78, CUDA2 2.33, CUDA3 2.42.

The same carve applies to the native GDN/QSA/shared-expert projections
(`NativeDense::load`'s range), which is 2,018.88 of those 3,485.66 MiB.

### Three things that fail silently

* **The lower bound.** The loader must be given `[lb, le)`, not `[0, le)`. `[0, le)` leaves two
  thirds of the hole unclaimed while every printed number looks plausible — this got as far as a
  commit before the per-range byte comparison caught it.
* **The routers stay on CUDA0.** CS-T's file-tier lookahead memcpy's `ffn_gate_inp` out of CUDA0's
  table to warm the next layer's pages. A dropped row is still in the metadata but has
  `data == nullptr`, so the copy fails, `ok` goes false, and the lookahead turns itself off
  **without saying so**. All 48 are kept — 120 MiB is cheaper than losing the prefetch.
* **`--split-skip-if-fits` measures CUDA0 before its weights exist**, so the full dense cost is
  added to what it holds back. Without that term the carve makes the check spend exactly the VRAM
  it is about to free.

`STRATA_WEIGHT_SLICE=0` is the kill switch, so one binary runs both A/B arms.
`STRATA_STAGE_CACHE_SLOTS=S` pins stage caches (test-only), because they are auto-sized from free
VRAM and freeing 2.5 GiB per stage would change the slot counts and mask the comparison.

## 2. Every card's leftover VRAM spills into a helper expert tier

The carve exposed a second problem it does not fix on its own: a stage's cache only ever holds **its
own layers'** pairs, so once a stage holds its whole range the cache cannot grow and the cost model
scores its spare VRAM at exactly **zero**. On the 4-way rig that stranded real memory — CUDA0 at
3,144 MiB and CUDA1 at 1,739 MiB while both 5060s were packed at 7,225 and 7,311.

With `--layer-split`, every card the run already uses — CUDA0 included — now hands its leftover to
a **helper tier** holding the globally hottest pairs no stage cache can hold. Stage misses hit it.

* **Automatic**, no new opt-in. `--no-remote-auto` / `STRATA_REMOTE_AUTO=0` is the kill switch, so
  it is revertable and A/B-able from the config JSON without a rebuild.
* Helpers **fail soft**: a tier that cannot open is logged and skipped. The explicit
  `--expert-cache-deviceN` path keeps its hard failure.
* Helpers stay **decode/verify-only** — they are reached from the pool dispatch, never from
  `Prefill`, whose GPU MoE path measured 2.2× slower and corrupting.
* The stage caches' own-layers rule is **unchanged**. Only the helper spans all layers, so the
  stage-relative `d_res` table is untouched.
* Tiers open *after* the placement is chosen and the caches are filled, so the search is
  **untouched** — it still picks what it picked.

The tier count goes 3 → 4 (`kRemoteMax`) to cover four cards, and a helper may now live on device 0
(`device < 1` → `device < 0`), which is why `preflight(0)` runs in the early window before anything
else creates that context.

Selection is by **byte budget**, not slot count: a slot count is a poor proxy for a native pack
where `blob_bytes(l)` varies per layer. The selection loop is extracted into a CUDA-free function
and unit-tested (`tests/core/remote_budget_test.cpp`, run under `CUDA_VISIBLE_DEVICES=-1`) for the
budget bound, `claimed`/primary-cache disjointness, rank order, and bit-for-bit agreement with the
old `slots` mode.

**Not yet measured on the rig.** The mechanism, the budget arithmetic and the disjointness are
covered by the unit test; the decode and hit-rate effect needs the 4-way run, which the cards were
busy for. Worth watching: the per-tier log line, and whether CUDA0's and CUDA1's stranded VRAM
shows up as tiers.

## 3. Score every four-way placement

The search tried every placement up to 3 GPUs and, beyond that, shared layers in proportion to
speed. Four cards got a heuristic instead of a search. It now enumerates all C(L−3, 3) = 15,180
placements and scores each. Memoised per-range fills (1,121 walks) bring it from 740 ms to 14 ms.

The mass model moved from `mass[r] = (r+1)^-1.2` to the increment of `H(f) = 1 − (1 − f)^b` over
the held fraction, `b = 3` (`STRATA_SPLIT_COVER_B`). At f = 0.8138 (20,000 held) it gives 99.35%,
identical to the old fit — so the validated 2-GPU regime is unchanged.

## 4. The Monitor's per-GPU cards

The split was invisible: one card holding 27 GiB read the same as four. Each GPU now gets its own
card, built from what only the engine knows (layer range, KV and prompt buffers, expert cache)
paired with that card's VRAM from NVML. The engine reports its allocation as seven parallel CSV
fields on the INFO line — single comma-separated tokens because the server splits on whitespace.
NVML's `mem_free` was being read and thrown away; each card's name is asked for once rather than
once a second. A device that both runs a stage and hosts a helper merges into one card rather than
appearing twice.

The image encoder is a separate process reporting no memory of its own, so its footprint is
measured as the card's free VRAM lost while it started — `strata-vision` runs a 2048px warm-up
before printing READY, and a delta outside 64–8192 MiB is discarded as noise, keeping the last good
reading. The encoder is pinned to the card the config's `vision.cuda_device` names (else the
engine's first), which is the card the Monitor must show it on.

## 5. The PLE table's Q8_0 row

`main` now takes the PLE table's row size from the tensor's own GGML type (IQ4_NL 90 B, Q5_0
110 B, FP8 E4M3 160 B). This branch adds the fourth format the published Q8_0 checkpoint carries:
**Q8_0, 170 B** — 5 blocks of 34, one fp16 `d` per 32 values, no min term. Same geometry, so the
row cache, the batch slab and the 16-rows-make-2560 flatten are untouched; only the stride and the
block decoder differ. `ple_row_bytes()`/`ple_dequant_row()` dispatch to the decoders in
`artifact/dequant.hpp`, so the PLE row and the artifact reader cannot drift apart on the block
layout. A row decoded at the wrong stride opens and decodes plausible garbage, so the row size and
format are logged once at open.

Unlike Q5_0, Q8_0 is **not** restricted to `--ple-io mmap`: the Direct reader takes the row size.

## 6. Three fixes found on the way

An image source that 404s, refuses the connection or names an unknown host raised `OSError`, which
the dispatcher does not map to a status — the client's connection closed with no reply at all
(`curl: (52) Empty reply from server`) instead of a 400. Those now raise `ValueError` like every
other bad image source. `free_vram_mib` is one helper taking an index rather than two copies of the
same NVML call.

## Testing

`tests/core/weights_carve_test.cpp` covers the range predicate and the priced-vs-loaded agreement.
`tests/core/remote_budget_test.cpp` covers the helper tier's selection.

There is no multi-GPU integration test in this repo, so the carve was verified by A/B: the engine
driven directly through `--serve` (no tokenizer, so stdin takes `GEN <n> <ids>` and stdout emits
`T <id>` per token), with an explicit split and everything pinned — explicit `--prefill`,
`STRATA_LOOKAHEAD=0`, `STRATA_STAGE_CACHE_SLOTS` — two binaries, one token hash, and the same hash
as before the change. The `--split-device 0` bit-exact A/B is only meaningful with
`--no-remote-auto`; with tiers on, tokens may legitimately differ from the CPU pool's float order.

## Merged `main`

`main` is merged into this branch (128 commits, 0.1.32–0.1.34) and the conflicts resolved instead of
re-opened: the PLE row-size design is `main`'s with Q8_0 re-added on top, and the upstream changes
in the split region (`Verifier::set_commit_async`, #340's streamed-ring reserve, `stage_room`'s
search/reserve distinction) are ported onto this branch's structure rather than replacing it.

Verified after the merge: builds clean, 56 of 58 ctest targets pass serially (the two failures are
environmental — `ple_parity` wants a model shard that is not in the tree, `expert_multi_test` wants
AVX512-VNNI/VBMI), and `serve`'s 120 Python tests pass.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

En el sitio

Enlaces a install, modelos, releases.