Issues / #1670
#1670 [Feature]: Layer-major prefill for streamed experts — 4–6× faster long prompts on a 6 GB card (PRs offered)
open · @vizdatom · 2 comments · View on GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quantsLinux
Description
### What do you want to do?
Make prompt reading fast on a small-VRAM card when the experts are streamed (`--mmap-experts`). On a 6 GB RTX 3060 Laptop with Qwen3.8-Flash-Next Q2_0, a layer-major prefill takes long prompts from **62–99 tok/s to 378–443 tok/s** (a 223.7K-token prompt from ~2,500 s to ~520 s), with decode about the same (19–21 tok/s). I have this working on a local branch and would like to open PRs if you're interested.
### Why it is slow today
With `--prefill auto` on 6 GB, the prompt is read in 512-token chunks and every chunk walks all layers, so each layer's routed experts are streamed from the file tier once per chunk. On this box a 12.9K-token prompt moves **~509 GB of expert bytes**, almost all of it the same experts re-read 25 times. #1407 looks like the same re-read pattern on a layer split ("file tier re-reads 24–35 GB per request").
### What I changed (all opt-in env switches, default off)
1. **`STRATA_PREFILL_LAYER_MAJOR=1` — layer-major prefill.** A super-chunk of up to `STRATA_PREFILL_LM_TOKENS` (default 8192, shrunk to free VRAM minus `STRATA_PREFILL_LM_MARGIN_MIB`) is read layer by layer. Each 512-token sub-chunk runs attention, the hyper-connection read, the router and the shared expert in order (GDN / QSA / PLE state advances as in the default path). The layer's routed experts then run once for the whole super-chunk: each expert is staged once, its rows from every sub-chunk go into one MMQ group, and the results are added per token. The super-chunk's residuals live in VRAM or pinned RAM (`STRATA_PREFILL_LM_R=vram|host|auto`, copies overlapped on the copy stream). Every streamed expert of a layer is prefetched when its attention starts, and the ones the routing skips are dropped (`STRATA_LM_PREFETCH=0` turns this off). Expert traffic drops by about S / 512: the 12.9K prompt goes from ~509 GB to ~79 GB.
- Two small kernels (`moe_shared_init`, `moe_scatter_add`).
- A stager fix the prefetch needs: a job now waits for the previous job on its pinned buffer, and `skip()` marks unneeded jobs ready without reading them. Before, two read threads could write the same buffer. The default walk consumes jobs in order, so it isn't affected.
- `on_chunk` still reports every sub-chunk; serve doesn't checkpoint mid-super-chunk.
- Not used for a layer split, a peer GPU, image prompts, the fused layout, or prompts under two chunks.
2. **`STRATA_PREFILL_DIRECT=1` — split page cache vs O_DIRECT by residency.** On the prompt path, each expert blob gets a `mincore` residency check. Mostly cached (≥ `STRATA_PREFILL_DIRECT_MIN`, default 50%) means memcpy from the mapping; otherwise it's read with O_DIRECT. Decode stays on the page cache. Prompts no longer evict the decode working set: prefill +3–27% and decode +11–17% on this box.
3. **`STRATA_PREFILL_WARM=1` / `STRATA_BIN_WARM=1`** — the per-layer walk warms this and the next layer's file-tier experts, and the pack's `experts.bin` is warmed too (Linux). Small: +1–9% prefill.
### Results
Ryzen 7 5800H, RTX 3060 Laptop 6 GB, 27 GiB RAM, Apacer AS2280P4 NVMe (Gen3, 256 GB), Ubuntu 24.04, driver 595.99.02.
Engine: 0.1.39 + 10 commits (`idx-int8-524k`), with your bit-plane Q2_0 kernel and Linux O_DIRECT tier commits cherry-picked.
Model: Qwen3.8-Flash-Next GSQ-RCO Q2_0 pack. Flags: `--mmap-experts --expert-cache auto --prefill auto --kv q4_0 --kv-resident 32768 --max-context 262144 --spec 3 --spec-min-p 0.5 --mtp … --vram-reserve-mib 650`.
Methodology: page cache evicted before each run, unique prefix per request (no reuse), temperature 0.6, thinking off; tok/s from the engine's own `strata serve: prompt …` log lines.
| | prompt 7.4K | prompt 8.1K | prompt 12.9K | prompt 51.5K | decode (prose / code) |
|---|---|---|---|---|---|
| 0.1.39 base | 69 | 74 | 95 | – | 17.6 / – |
| + bit-plane, both warms | 75 | 81 | 96 | – | 18.5 / 18.3 |
| + `PREFILL_DIRECT` | 95 | 98 | 99 | 62.5 | 20.6 / 21.4 |
| + `PREFILL_LAYER_MAJOR` | **378** | **406** | **443** | **415** | 19.1 / 19.8 |
**223,706-token needle prompt:** 2,400–2,611 s before, 498–538 s after (3/3 found).
**Correctness with layer-major on, served for agent use:**
- 17-check tool-calling test (tool choice from a menu, nested/optional/enum params, multi-turn chains, no-tool and error recovery): 17/17
- needle at 223.7K: 3/3
- BFCL `multi_turn_base`: 55.5% (200 entries, 7 h 52 min)
- engine errors: none across 2,547 requests
**Side finding:** with `--resident-budget-gib 14` on 27 GiB RAM, the Linux O_DIRECT tier picks unbuffered reads by itself and decode falls from 18.4 to 4.5 tok/s, because decode misses go to the NVMe. `STRATA_UNBUFFERED_LOAD=0` restores it. Prompt reads did get faster in that mode (117 vs 77 tok/s); item 2 above gets that gain without the decode loss.
### Proposal
My branch is 0.1.39-based and main has moved a lot (#1107, #1247, #1441 also touch prefill). Before rebasing, I'd like to know if you want this and in what shape. My suggestion is three PRs against current `main`, each re-measured after the rebase:
1. Layer-major prefill, with its kernels and the stager fix
2. `PREFILL_DIRECT`
3. The two warm switches
Happy to change names, defaults or structure to fit your plans for the prefill path.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.