Issues / #604

#604 Resident expert placement on a layer split, for mixed or older GPUs: where I think it helps, plus some doc notes, and an offer to test

closed · @paulhothersall · 4 comentários · No GitHub

BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quantsDocumentation

Descrição

Hi, and thanks for Strata. I'm posting because I started with one idea, found most of it was already built or planned, and want to offer hardware time for the part that isn't. Everything below is from **reading** `main` (99f3dbd), the docs, `bench/results/` and issues #364, #384, #395 and #498. **I have not run Strata yet and none of my numbers below are measurements.**

**Disclosure:** I used Claude (Anthropic's CLI) to read the repo and cross-check my reasoning. I checked the quoted file/line claims against `main` before posting. I'm not asking you to review a patch. There isn't one.

## Why I looked

I have older, cheaper cards: [Tesla P40s (I can run 1, 2 or 3 of them, 24 GB each), 64 GB host RAM, Ryzen  3900x, and sometimes a P40 beside a 3070. For this class of machine the question is: can **VRAM + RAM together** hold a bigger quant than RAM alone could, with each expert in one place instead of two?

## What I found is already there (credit where due)

- The resident modes already keep one copy: RAM holds only the experts no GPU holds. `--resident-experts` and `--resident-budget-gib` do this, and a swap copies the evicted expert back into the host slot of its replacement (`resident_stage_swaps` in `generate.cpp`, `stage_exchange` / `commit_exchanges` in `expert_source.cpp`).
- A hotness-ranked RAM budget with the rest on the SSD, a next-layer router lookahead that warms file pages, `--dump-routing`, and `--expert-profile-save`.
- The auto split's cost model, with a per-card speed term.
- Your reply on #384 that a real resident mode across a split "is on our list", and the 0.1.33 / 0.1.37 changes.
- The P40 work in #395 and the per-stage weight carve in #390.

So I'm **not** proposing a new swap design.

## The parts I think might still be useful

**1. Resident mode across a split, for predictable RAM.** From #384 the gain looks like RAM safety rather than speed: about 42 GB stayed available with `--mmap-experts` vs about 0.5 GB without it, at similar tokens/s. That matters most where the quant is bigger than RAM allows but VRAM + RAM is enough. My estimate for **P40-class 24 GB cards + 64 GB RAM** (formula calibrated against your 3090 IQ3_XXS run, which it reproduces within 2%; each card carries its own dense copy and session, so only about 17-18 GB of a 24 GB card is left for experts):

| Cards (24 GB each) | UD-Q4_K_XL (77 GB) in VRAM / in RAM | RAM needed incl. ~10 GB headroom | Today |
|---|---|---:|---|
| 1 | ~17 GB / ~60 GB | ~70 GB (more than 64) | budget mode + SSD tail (documented) |
| 2 | ~34 GB / ~43 GB | ~53 GB | setup keeps one GPU below ~135 GB RAM |
| 3 | ~51 GB / ~26 GB | ~36 GB | same |

IQ3_S (50 GB) already fits in 64 GB in the default mode, and on 3 cards it fits entirely in VRAM, so single-copy only changes anything for **Q4-class quants on 2-3 cards**. That is the case I'd aim at. Treat the numbers as estimates: the Q4 non-expert VRAM is back-solved from the doc's 1,280 slots on the 5070.

These per-card figures are **before #390** (the per-stage weight carve, which reclaims about 2.3-2.8 GiB per card in its measurements). With it, each card would hold more experts, so the RAM column would shrink a little more. I don't know how the two changes would interact; that's a question for you.

**2. Card speed in the split cost model.** The cost model in `generate.cpp` (`~2414`) fits per-layer time to SMs x clock on a 5080 + 3090. A mixed pair like P40 + 3070 differs in memory bandwidth, FP16/BF16 capability and PCIe, none of which that term sees. I don't know if it changes the chosen split in practice. That's a question for a test, not a claim.

**3. A small mechanical detail, only if measured.** Today a swap is D2H into an exchange buffer, a stream sync, H2D, then a CPU `memcpy` at the flip: about 4 blob-sized DRAM transfers, vs 2 with a pre-reserved pinned spare slot and two streams. I'd only mention it if `commit_exchanges` shows up in a profile. I don't have one.

## Things I noticed while reading (low-risk doc fixes, all re-checked on `main`)

1. `docs/MODELS.md:74` says RAM >= "shard 1 + about 10 GB". `setup.py:2030` uses experts + 10 (`arena_gb`). Shard 1 includes 3.6-4.5 GB of dense weights that live in VRAM.
2. `docs/DETAILS.md:890` says the n-gram table is read "through the OS cache". `direct_file.hpp` says it is unbuffered and must never use the file cache.
3. `docs/DETAILS.md:895` ("How it works") still says 2,048-token chunks; the same page documents 8,192 elsewhere.
4. `docs/DETAILS.md:280` says a "~40 GB" Q2_0 repack; `expert_source.hpp:378` gives `experts.bin` = 33,973,862,400 B. (The 40 may be a disk budget, in which case ignore this.)
5. The README speed table looks like it mixes versions: Q2_0 matches 0.1.36; the other rows match the 0.1.26 tables.
6. `expert_cache.hpp:19-25` still says the cache "does not yet compute anything"; `generate.cpp` calls that comment false.

## What I can run

I'm happy to be a test rig for anything useful here, especially the awkward hardware:

- **1, then 2, then 3 x Tesla P40 + 64 GB RAM:** the #395 branch (or experimental sm_61), single P40 first, then splits. With 3 identical cards I can also check whether the third card helps or just adds a per-window hand-off (your #384/#390 data suggests it can go either way).
- **P40 + 3070** (slow big card + fast small card): auto split vs forced splits, to see whether the speed term matters.
- Capture `--stats`, expert tiers, hit rates and `STRATA_TRACE_ADAPT` output, and run `--dump-routing` on 3 workloads (code, chat, random-text/needle). Your own data shows hit rate moves 15-20 points by workload at the same resident fraction.

**Questions:**
1. Is resident-on-split still wanted, and is there a design you'd prefer?
2. If I add a flag-off trace (pass id, tier, hit/miss, swaps, lookahead outcomes), where would you want it, or is `--dump-routing` enough?
3. Which P40 runs would be most useful to you?

I'm not claiming a speedup. Without measurements I only claim that a configuration some of us have becomes possible.

No site

Links install, modelos, releases.