Pull requests / #1050

#1050 prefill: STRATA_PREFILL_CPU_SHARE=auto hands a small chunk's least-routed experts to the idle CPU pool (opt-in)

closed · @sergqwer · 0 comentarios · En GitHub

Server & APINVIDIA / CUDAModels & quants

Descripción

A chunk below `stream_all_min()` (1,024 tokens) streams every routed expert outside VRAM over PCIe. An agent's tool result, a test's output and a short follow-up are all such chunks. Meanwhile the decode CPU pool sits idle, and the arena holds the same experts in RAM.

This adds an opt-in, `STRATA_PREFILL_CPU_SHARE=auto|x`. Unset (the default), no expert moves and the output is byte-identical to v0.1.39.

### What it does

- **The CPU's experts.** The pool takes the non-resident experts routed by at most `MAXT` of the chunk's tokens, fewest first. It runs them on a thread while the GPU streams and runs the rest.
- **Their rows.** The CPU's unweighted rows go to `Dm`'s tail, as a peer GPU's do (the same mechanism as `--peer-device`), and the combine weights them as any other row.
- **`auto`: the share is measured.** Each layer times the CPU thread per expert, and the GPU's expert work per streamed expert. The GPU side uses CUDA events around its expert work, read after the next layer's routing sync, because the host reaches the combine long before the GPU does; timing the host instead gave a wrong share. The next layers hand the CPU `g / (c + g)` of the streamed experts, the point where both sides end together. A small CPU therefore takes a small share and does not make the prompt slower. `x` fixes a share.
- **Activations.** Q2_0 layers take the pool's Q2_0 kernels' activations (`ActQ`, as decode's), and i-quant layers take the native `vec_dot` type.
- **Scope.** Serve without batch slots, and generate; one GPU.

### Measured

RTX 5090 + Ryzen 9 9950X3D, generate, one chunk, ms (off / auto):

| model | 600 tokens | 250 tokens |
|---|---|---|
| UD-Q4_K_XL (every expert in RAM) | 1,548 / 1,369, 1,528 / 1,296 (share 0.65) | |
| Q2_0 | 472 / 445, 491 / 448 (0.90) | |
| IQ2_XS | 493 / 481, 502 / 460, 527 / 529 (0.83-0.86) | 424 / 355, 402 / 382, 379 / 367 |
| IQ2_XS, 2 AVX2 pool workers (emulated) | 512 / 487 (0.49) | |

- **Where it helps most.** The gain grows with the experts' size. UD-Q4_K_XL's larger experts make a small chunk PCIe-bound. Q2_0's and IQ2_XS's ~1.2 MB experts mostly do not, so they move less.
- **Accuracy.** The CPU's rows are not bit-identical to the GPU's (another activation format), but they are as accurate.
  - IQ2_XS, first-token KL to `STRATA_PREFILL_MMQ=0`: 0.0053 off, 0.0080 auto.
  - UD-Q4_K_XL, KL between off and auto: 0.006.
  - On a fork with larger NVFP4 experts, layer 0's expert rows were 1.086% from the FP16 path on the CPU against 1.087% for the GPU's MMQ rows.
- **Identity.** Off, the logits and the 32 greedy tokens equal v0.1.39 on IQ2_XS at 600 and 2,100 tokens (fixed cache).
- **Not changed.** `stream_all_min()` stays 1,024. On the fork with NVFP4's 2.76 MB experts the routed-only walk paid up to 4,096 tokens, but on Q2_0 it lost against streaming every expert, so it stays out of this PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

En el sitio

Enlaces a install, modelos, releases.