Pull requests / #1282
#1282 prefill: STRATA_PREFILL_CPU_SHARE=auto hands a small chunk's least-routed experts to the idle CPU pool (opt-in)
closed · @sergqwer · 0 comentários · No GitHub
BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindows
Descrição
Replaces #1050, which GitHub closed when `main`'s history was rewritten. The closing note says this was not a rejection. It is the same commit, cherry-picked onto the new `main` (v0.1.40.1). --- A chunk below `stream_all_min()` (1,024 tokens) streams every routed expert outside VRAM over PCIe. An agent's tool result, a test's output and a short follow-up are all such chunks. Meanwhile the decode CPU pool sits idle, and the arena holds the same experts in RAM. This adds an opt-in, `STRATA_PREFILL_CPU_SHARE=auto|x`. Unset (the default), no expert moves and the output is byte-identical to v0.1.40.1. ### What it does - **The CPU's experts.** The pool takes the non-resident experts routed by at most `MAXT` of the chunk's tokens, fewest first. It runs them on a thread while the GPU streams and runs the rest. - **Their rows.** The CPU's unweighted rows go to `Dm`'s tail, as a peer GPU's do (the same mechanism as `--peer-device`), and the combine weights them as any other row. - **`auto`: the share is measured.** Each layer times the CPU thread per expert, and the GPU's expert work per streamed expert. The GPU side uses CUDA events around its expert work, read after the next layer's routing sync, because the host reaches the combine long before the GPU does; timing the host instead gave a wrong share. The next layers hand the CPU `g / (c + g)` of the streamed experts, the point where both sides end together. A small CPU therefore takes a small share and does not make the prompt slower. `x` fixes a share. - **Activations.** Q2_0 layers take the pool's Q2_0 kernels' activations (`ActQ`, as decode's), and i-quant layers take the native `vec_dot` type. - **Scope.** Serve without batch slots, and generate; one GPU. ### Measured **On another machine, by @did-technomancer (#1050):** RX 6800 16 GB (gfx1030) + Ryzen 7 5700X3D (AVX2, no AVX-512), Windows 11, ROCm, IQ3_XXS, 128K context. Short prompts (~850 tokens) read 37-45% faster, and long prompts and decode are unchanged. Their quality suite gave 22/22 with the switch on. | engine | short prompts, tok/s | prompt 8K / 30K, tok/s | |---|---|---| | v0.1.39 + this PR, off | 159-166 | 594 / 543 | | v0.1.39 + this PR, `auto` | **233-241** | 593 / 544 | | v0.1.40 + this PR, off | 176-186 | 739 / 761 | | v0.1.40 + this PR, `auto` | **241-255** | 735 / 758 | **On 0.1.40, here** (IQ2_XS, one chunk, ms, off / auto): 600 tokens 487 / 463, 489 / 457, 491 / 485 (share 0.86-0.87); 250 tokens 386 / 357, 386 / 352, 379 / 356 (share 0.90). **On 0.1.39, here.** RTX 5090 + Ryzen 9 9950X3D, generate, one chunk, ms (off / auto): | model | 600 tokens | 250 tokens | |---|---|---| | UD-Q4_K_XL (every expert in RAM) | 1,548 / 1,369, 1,528 / 1,296 (share 0.65) | | | Q2_0 | 472 / 445, 491 / 448 (0.90) | | | IQ2_XS | 493 / 481, 502 / 460, 527 / 529 (0.83-0.86) | 424 / 355, 402 / 382, 379 / 367 | | IQ2_XS, 2 AVX2 pool workers (emulated) | 512 / 487 (0.49) | | - **Where it helps most.** The gain grows with the experts' size. UD-Q4_K_XL's larger experts make a small chunk PCIe-bound. Q2_0's and IQ2_XS's ~1.2 MB experts mostly do not, so they move less. - **Accuracy.** The CPU's rows are not bit-identical to the GPU's (another activation format), but they are as accurate. - IQ2_XS, first-token KL to `STRATA_PREFILL_MMQ=0`: 0.0053 off, 0.0080 auto. - UD-Q4_K_XL, KL between off and auto: 0.006. - On a fork with larger NVFP4 experts, layer 0's expert rows were 1.086% from the FP16 path on the CPU against 1.087% for the GPU's MMQ rows. - **Identity.** Off, the logits and the 32 greedy tokens equal v0.1.40 on IQ2_XS after 600 and 32K tokens (same expert cache). 0.1.40.1 changed only the server and setup, and the engine's sources here are the same bytes as on that branch. - **Not changed.** `stream_all_min()` stays 1,024. On the fork with NVFP4's 2.76 MB experts the routed-only walk paid up to 4,096 tokens, but on Q2_0 it lost against streaming every expert, so it stays out of this PR. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
No site
Links install, modelos, releases.