Pull requests / #1414
#1414 prefill: the CPU share (STRATA_PREFILL_CPU_SHARE) up to 3,072-token chunks, from mapped experts, on a layer split
open · @architectds · 0 commentaires · Sur GitHub
Setup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Description
## Title
Issue: none - a follow-up to STRATA_PREFILL_CPU_SHARE (978d355, 0.1.40.2).
## Summary
`STRATA_PREFILL_CPU_SHARE` (0.1.40.2) hands the decode CPU pool a share of a small chunk's streamed experts. On a PC
that maps its experts (`--mmap-experts`) it took none, because it only takes page-locked experts; it also stopped at
1,024-token chunks and ran on one GPU only. With this change, when it is set:
- it takes any expert the CPU reads from RAM as it is: the arena, the page-locked copy, or the mapped `experts.bin`'s
pages in the file cache;
- chunks below 3,072 tokens are staged after their routing (only those can hand the CPU a share), instead of streaming
every expert from 1,024 tokens on; `STRATA_PREFILL_CPU_SHARE_MAX=1024` keeps the old limit;
- on a layer split every stage has the pool, and one stage at a time takes it for a chunk.
Unset, nothing changes: every new path is behind `cpu_share_on()`.
Measured on Windows 11, Ryzen 9 5900XT, 96 GB DDR4-3200, PCIe 3.0 x8, `--mmap-experts`, q4_0 KV. Fresh prompts read
after an 8K prewarm, medians of 5 interleaved rounds, off / auto, ms:
| machine | 512 tokens | 1,000 tokens | 2,000 tokens | 3,000 tokens |
|---|---|---|---|---|
| RTX 5070 Ti, Swift 1.5 IQ3_XXS (huihui-ai's build) | 3,322 / 1,706 (-49%) | 3,990 / 2,357 (-41%) | 5,413 / 3,354 (-38%) | 5,525 / 4,309 (-22%) |
| RTX 3060 + RTX 5070 Ti (layers 0-11 / 12-47), IQ3_S | 3,452 / 2,037 (-41%) | 4,506 / 2,680 (-41%) | 6,376 / 3,869 (-39%) | 6,733 / 5,138 (-24%) |
A coding agent's recorded conversation on the two cards (a 100K-token start, then 8 turns of 1-5K tokens of code, 128
tokens written a turn; both read the same tokens): 137.4 s off, 130.6 s auto (reading 121.1 -> 114.3 s).
## What changed
- `src/prefill/prefill.cpp`
- `stream_all_min()`: once the share is armed, the share's own limit (`STRATA_PREFILL_CPU_SHARE_MAX`, default 3072,
at least 1024) instead of 1024. `Prefill::arm_cpu_share` sets it before any loan is sized; `init` sets it too.
- The experts the share may take: page-locked, or not `transient()` with a stable blob. The blob pointers are taken
while the experts are chosen (on the prefill thread, as before); an expert without one stays on the GPU.
- A chunk takes the pool only when no other stage of a layer split holds it (`g_cpu_pool_busy`), and gives it back
at the end of the chunk, after the last layer's CPU thread has been joined.
- `include/strata/prefill/prefill.hpp`: `Prefill::arm_cpu_share`, and `set_cpu_pool`'s comment.
- `src/program/generate.cpp`
- serve: the limit is armed before the loans are sized, and every stage of a split gets the pool (still not with
`--batch`);
- the one-shot generate path: armed before `plan_lend` (one GPU, as before).
- `docs/DETAILS.md`: the CPU share section, with the table above.
## Extra Notes
- **Before this, on the same setup, the share took nothing.** With it on, the one-card short reads took off's times
(3,319 against 3,322 ms at 512 tokens; 3,000 tokens 5,487 against 5,525).
- **Output with it on changes in the last bits**, as the section already says. On one card, all 9 greedy answers of the
conversation differ from off, the first differing word coming after 4-41 words (near ties).
- **It has run for a while.** The fork architectds/Strata has had its own CPU share since 0.1.36, on by default in its
builds and in daily use for coding agents since 2026-10-02 (one RTX 5070 Ti, later with an RTX 3060 beside it). The
same tests measured it equal to this up to 1,000 tokens, and 3% / 5% faster at 2,000 / 3,000 tokens on one card (on
the two cards within 1.5%): it quantizes the layer's activations on the GPU and copies a quarter of the bytes down,
where this copies floats and quantizes on the CPU. That part can follow as its own PR.
- **How much it helps depends on the hardware.** It pays where streaming an expert costs more than computing it on the
CPU: a narrow or older PCIe link (3.0, x8 or x4, an eGPU, a laptop), a card whose cache holds few of the experts, and
a CPU with many cores and good RAM bandwidth - the PC above has all three (-41% to -49% at 1,000 tokens and below).
It helps less with PCIe 5.0 x16 or a card that holds most experts, and a slow CPU or low RAM bandwidth gets a small
share (`auto` measures both sides). The table in DETAILS.md already has -19% to -35% below 1,000 tokens on three other
PCs. Not measured here: PCIe 4.0 / 5.0 links, AMD or Intel GPUs, Linux.
- **Tested:** Windows, CUDA 13.0, sm_86 + sm_120 (RTX 3060 + RTX 5070 Ti). Not tested: Linux, HIP, SYCL.
Sur le site
Liens install, modèles, releases.