Pull requests / #390
#390 Multi-GPU: a --layer-split stage holds only its own layers' weights, and every card's leftover VRAM spills into a helper expert tier (4-way IQ3_S: 9,269 -> 14,172 expert slots, prefill chunk 2,048 -> 4,096)
closed · @gopinath87607 · 0 comentarios · En GitHub
Server & APIMulti-GPUNVIDIA / CUDAModels & quants
Descripción
## What this is Three things a `--layer-split` run needed, plus the Monitor work that made the first one visible. The headline is the **carve**: every stage used to load the whole model's dense weights. Then the VRAM those weights were wasting is given back — first to each stage's own expert cache, and then, for whatever is still left over, to a **helper tier** that any stage's misses can hit. ## 1. A stage holds only its own layers' weights Each `--layer-split` stage loaded the **whole model's** dense weights: on the 4-GPU IQ3_S rig, 3,485.66 MiB per card — 1,466.78 pack arena + 2,018.88 native dense — about 10.2 GiB across four cards, holding layers that stage will never run. A stage running 9 of 48 layers still held all 48. Weights now get the treatment the session already got. `session_bytes(g, cells, k, lo, hi)` is pure arithmetic, so the split search *prices* a candidate layer range before allocating it; weights were the one per-stage resource that never got that. The carve extends the same rule: a layer-range predicate on the loader's existing skip/compact path, and the per-stage load moved down to after `st.lb/le` are set. One rule produces both the priced size and the loaded layout, so they cannot disagree. Measured, 4-way IQ3_S (split 10,19,34 auto): | | before | after | |---|---|---| | expert slots | 9,269 | **14,172** | | expert cache | 17,388 MiB | **26,892 MiB** | | CUDA0 slots | 3,657 | **5,120** | | auto prefill chunk | 2,048 | **4,096** | Per-card reclaim: CUDA0 2.56 GiB, CUDA1 2.78, CUDA2 2.33, CUDA3 2.42. The same carve applies to the native GDN/QSA/shared-expert projections (`NativeDense::load`'s range), which is 2,018.88 of those 3,485.66 MiB. ### Three things that fail silently * **The lower bound.** The loader must be given `[lb, le)`, not `[0, le)`. `[0, le)` leaves two thirds of the hole unclaimed while every printed number looks plausible — this got as far as a commit before the per-range byte comparison caught it. * **The routers stay on CUDA0.** CS-T's file-tier lookahead memcpy's `ffn_gate_inp` out of CUDA0's table to warm the next layer's pages. A dropped row is still in the metadata but has `data == nullptr`, so the copy fails, `ok` goes false, and the lookahead turns itself off **without saying so**. All 48 are kept — 120 MiB is cheaper than losing the prefetch. * **`--split-skip-if-fits` measures CUDA0 before its weights exist**, so the full dense cost is added to what it holds back. Without that term the carve makes the check spend exactly the VRAM it is about to free. `STRATA_WEIGHT_SLICE=0` is the kill switch, so one binary runs both A/B arms. `STRATA_STAGE_CACHE_SLOTS=S` pins stage caches (test-only), because they are auto-sized from free VRAM and freeing 2.5 GiB per stage would change the slot counts and mask the comparison. ## 2. Every card's leftover VRAM spills into a helper expert tier The carve exposed a second problem it does not fix on its own: a stage's cache only ever holds **its own layers'** pairs, so once a stage holds its whole range the cache cannot grow and the cost model scores its spare VRAM at exactly **zero**. On the 4-way rig that stranded real memory — CUDA0 at 3,144 MiB and CUDA1 at 1,739 MiB while both 5060s were packed at 7,225 and 7,311. With `--layer-split`, every card the run already uses — CUDA0 included — now hands its leftover to a **helper tier** holding the globally hottest pairs no stage cache can hold. Stage misses hit it. * **Automatic**, no new opt-in. `--no-remote-auto` / `STRATA_REMOTE_AUTO=0` is the kill switch, so it is revertable and A/B-able from the config JSON without a rebuild. * Helpers **fail soft**: a tier that cannot open is logged and skipped. The explicit `--expert-cache-deviceN` path keeps its hard failure. * Helpers stay **decode/verify-only** — they are reached from the pool dispatch, never from `Prefill`, whose GPU MoE path measured 2.2× slower and corrupting. * The stage caches' own-layers rule is **unchanged**. Only the helper spans all layers, so the stage-relative `d_res` table is untouched. * Tiers open *after* the placement is chosen and the caches are filled, so the search is **untouched** — it still picks what it picked. The tier count goes 3 → 4 (`kRemoteMax`) to cover four cards, and a helper may now live on device 0 (`device < 1` → `device < 0`), which is why `preflight(0)` runs in the early window before anything else creates that context. Selection is by **byte budget**, not slot count: a slot count is a poor proxy for a native pack where `blob_bytes(l)` varies per layer. The selection loop is extracted into a CUDA-free function and unit-tested (`tests/core/remote_budget_test.cpp`, run under `CUDA_VISIBLE_DEVICES=-1`) for the budget bound, `claimed`/primary-cache disjointness, rank order, and bit-for-bit agreement with the old `slots` mode. **Not yet measured on the rig.** The mechanism, the budget arithmetic and the disjointness are covered by the unit test; the decode and hit-rate effect needs the 4-way run, which the cards were busy for. Worth watching: the per-tier log line, and whether CUDA0's and CUDA1's stranded VRAM shows up as tiers. ## 3. Score every four-way placement The search tried every placement up to 3 GPUs and, beyond that, shared layers in proportion to speed. Four cards got a heuristic instead of a search. It now enumerates all C(L−3, 3) = 15,180 placements and scores each. Memoised per-range fills (1,121 walks) bring it from 740 ms to 14 ms. The mass model moved from `mass[r] = (r+1)^-1.2` to the increment of `H(f) = 1 − (1 − f)^b` over the held fraction, `b = 3` (`STRATA_SPLIT_COVER_B`). At f = 0.8138 (20,000 held) it gives 99.35%, identical to the old fit — so the validated 2-GPU regime is unchanged. ## 4. The Monitor's per-GPU cards The split was invisible: one card holding 27 GiB read the same as four. Each GPU now gets its own card, built from what only the engine knows (layer range, KV and prompt buffers, expert cache) paired with that card's VRAM from NVML. The engine reports its allocation as seven parallel CSV fields on the INFO line — single comma-separated tokens because the server splits on whitespace. NVML's `mem_free` was being read and thrown away; each card's name is asked for once rather than once a second. A device that both runs a stage and hosts a helper merges into one card rather than appearing twice. The image encoder is a separate process reporting no memory of its own, so its footprint is measured as the card's free VRAM lost while it started — `strata-vision` runs a 2048px warm-up before printing READY, and a delta outside 64–8192 MiB is discarded as noise, keeping the last good reading. The encoder is pinned to the card the config's `vision.cuda_device` names (else the engine's first), which is the card the Monitor must show it on. ## 5. The PLE table's Q8_0 row `main` now takes the PLE table's row size from the tensor's own GGML type (IQ4_NL 90 B, Q5_0 110 B, FP8 E4M3 160 B). This branch adds the fourth format the published Q8_0 checkpoint carries: **Q8_0, 170 B** — 5 blocks of 34, one fp16 `d` per 32 values, no min term. Same geometry, so the row cache, the batch slab and the 16-rows-make-2560 flatten are untouched; only the stride and the block decoder differ. `ple_row_bytes()`/`ple_dequant_row()` dispatch to the decoders in `artifact/dequant.hpp`, so the PLE row and the artifact reader cannot drift apart on the block layout. A row decoded at the wrong stride opens and decodes plausible garbage, so the row size and format are logged once at open. Unlike Q5_0, Q8_0 is **not** restricted to `--ple-io mmap`: the Direct reader takes the row size. ## 6. Three fixes found on the way An image source that 404s, refuses the connection or names an unknown host raised `OSError`, which the dispatcher does not map to a status — the client's connection closed with no reply at all (`curl: (52) Empty reply from server`) instead of a 400. Those now raise `ValueError` like every other bad image source. `free_vram_mib` is one helper taking an index rather than two copies of the same NVML call. ## Testing `tests/core/weights_carve_test.cpp` covers the range predicate and the priced-vs-loaded agreement. `tests/core/remote_budget_test.cpp` covers the helper tier's selection. There is no multi-GPU integration test in this repo, so the carve was verified by A/B: the engine driven directly through `--serve` (no tokenizer, so stdin takes `GEN <n> <ids>` and stdout emits `T <id>` per token), with an explicit split and everything pinned — explicit `--prefill`, `STRATA_LOOKAHEAD=0`, `STRATA_STAGE_CACHE_SLOTS` — two binaries, one token hash, and the same hash as before the change. The `--split-device 0` bit-exact A/B is only meaningful with `--no-remote-auto`; with tiers on, tokens may legitimately differ from the CPU pool's float order. ## Merged `main` `main` is merged into this branch (128 commits, 0.1.32–0.1.34) and the conflicts resolved instead of re-opened: the PLE row-size design is `main`'s with Q8_0 re-added on top, and the upstream changes in the split region (`Verifier::set_commit_async`, #340's streamed-ring reserve, `stage_room`'s search/reserve distinction) are ported onto this branch's structure rather than replacing it. Verified after the merge: builds clean, 56 of 58 ctest targets pass serially (the two failures are environmental — `ple_parity` wants a model shard that is not in the tree, `expert_multi_test` wants AVX512-VNNI/VBMI), and `serve`'s 120 Python tests pass. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
En el sitio
Enlaces a install, modelos, releases.