Pull requests / #1416

#1416 prefill: the CPU share's activations quantized on the GPU, copied down beside the routing sync (builds on #1414)

open · @architectds · 0 comentarios · En GitHub

AMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

Descripción

## Title
Issue: none - builds on #1414 (this branch carries its two commits; the new one is the last, 6848915).

## Summary
With `STRATA_PREFILL_CPU_SHARE` on, every layer of a staged chunk copied its MoE input to the host as floats (T x n_embd
x 4 bytes) before the layer's routing sync, and the share's thread then quantized the tokens the CPU's experts read.
Now a layer whose gate/up take Q8_K or Q8_0 activations has them quantized on the GPU (`quantize_q8_K` /
`quantize_q8_0`, ggml's bytes) into a small device buffer, which is copied down on a stream of its own: about a quarter
of the bytes, and the routing sync no longer waits for them. The thread waits for that copy instead of quantizing.
Q2_0 gate/up layers (ActQ activations) copy the floats as before.

Measured against #1414 in the same window (Windows 11, Ryzen 9 5900XT, 96 GB DDR4-3200, PCIe 3.0 x8, `--mmap-experts`,
q4_0 KV): fresh prompts after an 8K prewarm, medians of 5 interleaved rounds, #1414 / this, ms:

| machine | 512 tokens | 1,000 tokens | 2,000 tokens | 3,000 tokens |
|---|---|---|---|---|
| RTX 5070 Ti, Swift 1.5 IQ3_XXS (huihui-ai's build) | 1,683 / 1,715 | 2,318 / 2,326 | 3,513 / 3,435 (-2.2%) | 4,152 / 4,005 (-3.5%) |
| RTX 3060 + RTX 5070 Ti (layers 0-11 / 12-47), IQ3_S | 2,026 / 2,039 | 2,677 / 2,645 | 4,081 / 3,909 (-4.2%) | 5,242 / 5,048 (-3.7%) |

Up to 1,000 tokens the two are the same within noise; the larger the chunk, the more floats it no longer copies. The
same arms run in two windows two hours apart differed by up to 4%, so these are small gains, measured on one PC.

## What changed
- `src/prefill/prefill.cpp`
  - `Impl`: `cpu_qdev` / `cpu_qhost` (T x `kNativeActBytes`, on the device and page-locked; at most 12 MB of VRAM at a
    3,071-token chunk, allocated on first use), a stream `cpu_cs` and two events; `release` frees them.
  - Per layer with the share: `mixed` is quantized on the main stream once the previous layer's copy has finished
    reading `cpu_qdev`, and T x `act_bytes` are copied down on `cpu_cs`. The CPU thread waits for that copy and points
    its jobs at those bytes.
  - When the buffers cannot be had, a line says so once and every layer copies the floats as before.

## Extra Notes
- **Output:** the GPU's `quantize_q8_K` is ggml's `quantize_row_q8_K_ref` transcribed (with its parity test), and
  `quantize_q8_0` writes ggml's bytes. So the CPU's experts should read the same bytes `native_quant_act` gave them,
  and the output should match #1414's. I did not compare outputs here.
- **Origin:** this is how the fork architectds/Strata's own CPU share has moved its activations since 2026-10-01. The fork
  measured its own share 3% / 5% ahead of #1414 at 2,000 / 3,000 tokens on one card; this closes most of that.
- **Tested:** Windows, CUDA 13.0, sm_86 + sm_120 (RTX 3060 + RTX 5070 Ti). Not tested: Linux, HIP, SYCL.

En el sitio

Enlaces a install, modelos, releases.