Pull requests / #1416
#1416 prefill: the CPU share's activations quantized on the GPU, copied down beside the routing sync (builds on #1414)
open · @architectds · 0 comments · View on GitHub
AMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux
Description
## Title
Issue: none - builds on #1414 (this branch carries its two commits; the new one is the last, 6848915).
## Summary
With `STRATA_PREFILL_CPU_SHARE` on, every layer of a staged chunk copied its MoE input to the host as floats (T x n_embd
x 4 bytes) before the layer's routing sync, and the share's thread then quantized the tokens the CPU's experts read.
Now a layer whose gate/up take Q8_K or Q8_0 activations has them quantized on the GPU (`quantize_q8_K` /
`quantize_q8_0`, ggml's bytes) into a small device buffer, which is copied down on a stream of its own: about a quarter
of the bytes, and the routing sync no longer waits for them. The thread waits for that copy instead of quantizing.
Q2_0 gate/up layers (ActQ activations) copy the floats as before.
Measured against #1414 in the same window (Windows 11, Ryzen 9 5900XT, 96 GB DDR4-3200, PCIe 3.0 x8, `--mmap-experts`,
q4_0 KV): fresh prompts after an 8K prewarm, medians of 5 interleaved rounds, #1414 / this, ms:
| machine | 512 tokens | 1,000 tokens | 2,000 tokens | 3,000 tokens |
|---|---|---|---|---|
| RTX 5070 Ti, Swift 1.5 IQ3_XXS (huihui-ai's build) | 1,683 / 1,715 | 2,318 / 2,326 | 3,513 / 3,435 (-2.2%) | 4,152 / 4,005 (-3.5%) |
| RTX 3060 + RTX 5070 Ti (layers 0-11 / 12-47), IQ3_S | 2,026 / 2,039 | 2,677 / 2,645 | 4,081 / 3,909 (-4.2%) | 5,242 / 5,048 (-3.7%) |
Up to 1,000 tokens the two are the same within noise; the larger the chunk, the more floats it no longer copies. The
same arms run in two windows two hours apart differed by up to 4%, so these are small gains, measured on one PC.
## What changed
- `src/prefill/prefill.cpp`
- `Impl`: `cpu_qdev` / `cpu_qhost` (T x `kNativeActBytes`, on the device and page-locked; at most 12 MB of VRAM at a
3,071-token chunk, allocated on first use), a stream `cpu_cs` and two events; `release` frees them.
- Per layer with the share: `mixed` is quantized on the main stream once the previous layer's copy has finished
reading `cpu_qdev`, and T x `act_bytes` are copied down on `cpu_cs`. The CPU thread waits for that copy and points
its jobs at those bytes.
- When the buffers cannot be had, a line says so once and every layer copies the floats as before.
## Extra Notes
- **Output:** the GPU's `quantize_q8_K` is ggml's `quantize_row_q8_K_ref` transcribed (with its parity test), and
`quantize_q8_0` writes ggml's bytes. So the CPU's experts should read the same bytes `native_quant_act` gave them,
and the output should match #1414's. I did not compare outputs here.
- **Origin:** this is how the fork architectds/Strata's own CPU share has moved its activations since 2026-10-01. The fork
measured its own share 3% / 5% ahead of #1414 at 2,000 / 3,000 tokens on one card; this closes most of that.
- **Tested:** Windows, CUDA 13.0, sm_86 + sm_120 (RTX 3060 + RTX 5070 Ti). Not tested: Linux, HIP, SYCL.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.