Issues / #1670
#1670 [Feature]: Layer-major prefill for streamed experts — 4–6× faster long prompts on a 6 GB card (PRs offered)
open · @vizdatom · 2 comentários · No GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quantsLinux
Descrição
### What do you want to do?
Make prompt reading fast on a small-VRAM card when the experts are streamed (`--mmap-experts`). On a 6 GB RTX 3060 Laptop with Qwen3.8-Flash-Next Q2_0, a layer-major prefill takes long prompts from **62–99 tok/s to 378–443 tok/s** (a 223.7K-token prompt from ~2,500 s to ~520 s), with decode about the same (19–21 tok/s). I have this working on a local branch and would like to open PRs if you're interested.
### Why it is slow today
With `--prefill auto` on 6 GB, the prompt is read in 512-token chunks and every chunk walks all layers, so each layer's routed experts are streamed from the file tier once per chunk. On this box a 12.9K-token prompt moves **~509 GB of expert bytes**, almost all of it the same experts re-read 25 times. #1407 looks like the same re-read pattern on a layer split ("file tier re-reads 24–35 GB per request").
### What I changed (all opt-in env switches, default off)
1. **`STRATA_PREFILL_LAYER_MAJOR=1` — layer-major prefill.** A super-chunk of up to `STRATA_PREFILL_LM_TOKENS` (default 8192, shrunk to free VRAM minus `STRATA_PREFILL_LM_MARGIN_MIB`) is read layer by layer. Each 512-token sub-chunk runs attention, the hyper-connection read, the router and the shared expert in order (GDN / QSA / PLE state advances as in the default path). The layer's routed experts then run once for the whole super-chunk: each expert is staged once, its rows from every sub-chunk go into one MMQ group, and the results are added per token. The super-chunk's residuals live in VRAM or pinned RAM (`STRATA_PREFILL_LM_R=vram|host|auto`, copies overlapped on the copy stream). Every streamed expert of a layer is prefetched when its attention starts, and the ones the routing skips are dropped (`STRATA_LM_PREFETCH=0` turns this off). Expert traffic drops by about S / 512: the 12.9K prompt goes from ~509 GB to ~79 GB.
- Two small kernels (`moe_shared_init`, `moe_scatter_add`).
- A stager fix the prefetch needs: a job now waits for the previous job on its pinned buffer, and `skip()` marks unneeded jobs ready without reading them. Before, two read threads could write the same buffer. The default walk consumes jobs in order, so it isn't affected.
- `on_chunk` still reports every sub-chunk; serve doesn't checkpoint mid-super-chunk.
- Not used for a layer split, a peer GPU, image prompts, the fused layout, or prompts under two chunks.
2. **`STRATA_PREFILL_DIRECT=1` — split page cache vs O_DIRECT by residency.** On the prompt path, each expert blob gets a `mincore` residency check. Mostly cached (≥ `STRATA_PREFILL_DIRECT_MIN`, default 50%) means memcpy from the mapping; otherwise it's read with O_DIRECT. Decode stays on the page cache. Prompts no longer evict the decode working set: prefill +3–27% and decode +11–17% on this box.
3. **`STRATA_PREFILL_WARM=1` / `STRATA_BIN_WARM=1`** — the per-layer walk warms this and the next layer's file-tier experts, and the pack's `experts.bin` is warmed too (Linux). Small: +1–9% prefill.
### Results
Ryzen 7 5800H, RTX 3060 Laptop 6 GB, 27 GiB RAM, Apacer AS2280P4 NVMe (Gen3, 256 GB), Ubuntu 24.04, driver 595.99.02.
Engine: 0.1.39 + 10 commits (`idx-int8-524k`), with your bit-plane Q2_0 kernel and Linux O_DIRECT tier commits cherry-picked.
Model: Qwen3.8-Flash-Next GSQ-RCO Q2_0 pack. Flags: `--mmap-experts --expert-cache auto --prefill auto --kv q4_0 --kv-resident 32768 --max-context 262144 --spec 3 --spec-min-p 0.5 --mtp … --vram-reserve-mib 650`.
Methodology: page cache evicted before each run, unique prefix per request (no reuse), temperature 0.6, thinking off; tok/s from the engine's own `strata serve: prompt …` log lines.
| | prompt 7.4K | prompt 8.1K | prompt 12.9K | prompt 51.5K | decode (prose / code) |
|---|---|---|---|---|---|
| 0.1.39 base | 69 | 74 | 95 | – | 17.6 / – |
| + bit-plane, both warms | 75 | 81 | 96 | – | 18.5 / 18.3 |
| + `PREFILL_DIRECT` | 95 | 98 | 99 | 62.5 | 20.6 / 21.4 |
| + `PREFILL_LAYER_MAJOR` | **378** | **406** | **443** | **415** | 19.1 / 19.8 |
**223,706-token needle prompt:** 2,400–2,611 s before, 498–538 s after (3/3 found).
**Correctness with layer-major on, served for agent use:**
- 17-check tool-calling test (tool choice from a menu, nested/optional/enum params, multi-turn chains, no-tool and error recovery): 17/17
- needle at 223.7K: 3/3
- BFCL `multi_turn_base`: 55.5% (200 entries, 7 h 52 min)
- engine errors: none across 2,547 requests
**Side finding:** with `--resident-budget-gib 14` on 27 GiB RAM, the Linux O_DIRECT tier picks unbuffered reads by itself and decode falls from 18.4 to 4.5 tok/s, because decode misses go to the NVMe. `STRATA_UNBUFFERED_LOAD=0` restores it. Prompt reads did get faster in that mode (117 vs 77 tok/s); item 2 above gets that gain without the decode loss.
### Proposal
My branch is 0.1.39-based and main has moved a lot (#1107, #1247, #1441 also touch prefill). Before rebasing, I'd like to know if you want this and in what shape. My suggestion is three PRs against current `main`, each re-measured after the rebase:
1. Layer-major prefill, with its kernels and the stager fix
2. `PREFILL_DIRECT`
3. The two warm switches
Happy to change names, defaults or structure to fit your plans for the prefill path.
No site
Links install, modelos, releases.