Pull requests / #580
#580 Weight carve 0138
closed · @gopinath87607 · 0 comentarios · En GitHub
Server & APIMulti-GPUNVIDIA / CUDAModels & quants
Descripción
## What this is
A `--layer-split` stage loaded the **whole model's** dense weights. On the 4-GPU IQ3_S rig that is
3,485.66 MiB per card — 1,466.78 pack arena + 2,018.88 native dense — about 10.2 GiB across four
cards, holding layers that stage will never run. A stage running 9 of 48 layers still held all 48.
Each stage now holds only its own `[lb, le)` layers' weights. The VRAM that frees goes to that
stage's expert cache.
This is the first of the pieces #390 was cut into. Nothing else from that branch is in here: no
placement search, no helper tier, no Monitor change, no image-API change. The split behaviour is
**identical to before** — the carve changes where the weights live, not what the run computes.
## What changed
Weights now get the treatment the session already got. `session_bytes(g, cells, k, lo, hi)` is pure
arithmetic, so the split search *prices* a candidate layer range before allocating it; weights were
the one per-stage resource that never got that. The carve extends the same rule: a layer-range
predicate on the loader's existing skip/compact path, and the per-stage load moved down to after
`st.lb/le` are set. One rule produces both the priced size and the loaded layout, so they cannot
disagree.
The ranges reach:
* `WeightPool::load(..., layer_lo, layer_hi)` — the pack arena (`src/core/weights.cpp`)
* `WeightPool::pool_bytes(..., layer_lo, layer_hi)` — the priced size, same predicate
* `NativeDense::load(..., layer_lo, layer_hi)` — the GDN/QSA/shared-expert projections, 2,018.88 of
those 3,485.66 MiB (`src/core/native_dense.cpp`)
* `NativeDense::served_layer_bytes()` — the priced size for the native half
* the per-stage load in `src/program/generate.cpp`, moved after the split is chosen
* CUDA0's own load, which takes the same range so the first card is not the one left holding
everything
## Three things that fail silently
* **The lower bound.** The loader must be given `[lb, le)`, not `[0, le)`. `[0, le)` leaves two
thirds of the hole unclaimed while every printed number looks plausible — this got as far as a
commit before the per-range byte comparison caught it.
* **The routers stay on CUDA0.** CS-T's file-tier lookahead memcpy's `ffn_gate_inp` out of CUDA0's
table to warm the next layer's pages. A dropped row is still in the metadata but has
`data == nullptr`, so the copy fails, `ok` goes false, and the lookahead turns itself off
**without saying so**. All 48 are kept — 120 MiB is cheaper than losing the prefetch.
* **`--split-skip-if-fits` measures CUDA0 before its weights exist**, so the full dense cost is
added to what it holds back. Without that term the carve makes the check spend exactly the VRAM
it is about to free.
## What it buys
The rig's own 4-way IQ3_S config, nothing pinned — `--layer-split auto`, `--expert-cache auto`,
`--prefill auto:32768`, `--kv int8 --kv-resident 32768`, context 262144, one run of each arm:
| | before | after |
|---|---|---|
| expert slots | 8,790 | **14,061** |
| expert cache | 17,372 MiB | **27,727 MiB** |
| CUDA0 slots | 1,549 | **2,974** |
| auto prefill chunk | 2,048 | **6,144** |
Both arms chose the **same** split, 0-13 / 14-28 / 29-37 / 38-47, so the VRAM is the only thing that
moved. Per card:
| stage | dense before | dense after | reclaimed | its cache |
|---|---|---|---|---|
| CUDA0 [0,14) | 3,485.66 MiB | — | 2,514 MiB | 2,723 → 5,237 MiB |
| CUDA1 [14,29) | 3,485.66 MiB | 1,095.30 MiB | 2,390 MiB | 2,572 → 5,111 MiB |
| CUDA2 [29,38) | 3,485.66 MiB | 661.06 MiB | 2,825 MiB | 6,594 → 8,960 MiB |
| CUDA3 [38,48) | 3,485.66 MiB | 720.82 MiB | 2,765 MiB | 5,489 → 8,427 MiB |
The same carve applies to the native GDN/QSA/shared-expert projections
(`NativeDense::load`'s range), which is 2,018.88 of those 3,485.66 MiB.
## Testing
`tests/core/weights_carve_test.cpp` covers the range predicate and the priced-vs-loaded agreement:
every tensor the loader keeps is in range, every tensor it drops is out of range, and
`pool_bytes(lo, hi)` equals the bytes `load(lo, hi)` actually takes — over every range, on a
synthetic pack built to exercise the rounding, and on the real IQ3_S pack when one is named:
./build/weights_carve_test /home/gopi/Strata-data/packs/iq3_s
weights_carve_test: OK (synthetic; and the real pack at .../packs/iq3_s)
It also pins the no-op: a full range with no `skip` reproduces the pack's own pool byte for byte.
There is no multi-GPU integration test in this repo, so the equivalence claim is an A/B rather than
a unit test. `STRATA_WEIGHT_SLICE=0` runs the pre-carve path and `=1` the carve, from **one binary**, so
the two arms differ in nothing but where the weights live. Everything else is pinned, because an
auto-sized cache would read the freed VRAM and change the slot counts — and then a token difference
would say nothing about the weights:
* an explicit `--layer-split`, so the placement is not a variable
* `--expert-cache` small enough that CUDA0's cache is not clamped differently in the two arms
* `--vram-reserve-mib` explicit, so the adaptive reserve stays out
* `STRATA_STAGE_CACHE_SLOTS=S` pins every stage's cache to the same slot count in both arms
* `--prefill` explicit, `STRATA_LOOKAHEAD=0`
The two arms, 4-way IQ3_S, `--layer-split 10,19,34`, `--expert-cache 512`,
`STRATA_STAGE_CACHE_SLOTS=512`, context 16384, a 512-token prompt and 32 greedy tokens:
| | `STRATA_WEIGHT_SLICE=0` | `=1` |
|---|---|---|
| CUDA1 dense | 1,466.78 + 2,018.88 MiB | **272.83 + 364.85 MiB** |
| CUDA2 dense | 1,466.78 + 2,018.88 MiB | **448.02 + 656.07 MiB** |
| CUDA3 dense | 1,466.78 + 2,018.88 MiB | **419.66 + 588.55 MiB** |
| stage caches | 512 / 512 / 512 slots | 512 / 512 / 512 slots |
| CUDA0 cache | 713 slots, 1.27 GiB | 713 slots, 1.27 GiB |
| `expert_slots` | 2,249 | 2,249 |
| `vram_free_mib` | 1,669 | **4,463** |
| **token ids** | `3680a63c…` | `3680a63c…` |
Same slot counts, same placement, **the same 32 token ids**, byte for byte — and 2,794 MiB more free
VRAM on CUDA0. That is the carve moving bytes and nothing else.
The ids are the 32 `T` lines each arm's log prints, in order, one per line:
grep -oE '^T -?[0-9]+' log | awk '{print $2}' | sha256sum
3680a63ccf51f88776d66b6b5a7c346fff291d42ed1ede6b50639d588387d6eb5
`STRATA_WEIGHT_SLICE=0` and `STRATA_STAGE_CACHE_SLOTS` are test hooks in the same spirit as the
existing `STRATA_TEST_CACHE_FAIL`; the first is also the revert switch, so the carve can be A/B'd
from a config without a rebuild.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
En el sitio
Enlaces a install, modelos, releases.