Pull requests / #1033
#1033 prefill: gather GPU-resident experts in groups on short prompts too
closed · @brenoperucchi · 0 comentarios · En GitHub
BenchmarksSetup & installNVIDIA / CUDAModels & quantsWindows
Descripción
#372 (in 0.1.38) gathers an MMQ group's experts in one launch, but only in the streamed walk. A chunk below `stream_all_min()` (1024 tokens) still takes the staged walk, where every expert gets its own `gather_native` launch. @BlueKingMuch noticed this in #372: the 923-token tail of a 64K prompt was not affected. On our box most short requests are 150-600 tokens with almost every routed expert already in the GPU cache, so this per-expert launch is where their time goes. With `STRATA_PREFILL_TIMING=1`, a ~450-token prompt spends 50-55% of its GPU timeline in the "dequant" phase, which on the MMQ path is the gather. The staged walk was left out of #372 because a staged expert sits in the 8-slot staging ring, which is smaller than a group of 16, so a group cannot hold its slots. Resident experts have no ring slot, no wait and no release. This patch groups only those: - a resident expert joins the open group (`gather_native_group`, the same bytes into the same group slots); - a staged expert first flushes what the group holds, then is gathered alone as before, and the group restarts after it; - products, their order, the MMQ tail memsets and the streamed walk are unchanged; - `STRATA_PREFILL_GROUP_GATHER=0` still turns grouping off everywhere. One file, `src/prefill/prefill.cpp`, +17/-2. Tested on an RTX 5090 32 GB, Ryzen 9 5950X, 96 GB DDR4, Windows 11 (WDDM), CUDA 13.0, with the Swift IQ3_XXS native pack, `--expert-cache auto`, `--prefill auto:32768`, `--kv int8`, `--max-context 32768`. My builds use the setup.py flags plus `STRATA_PORTABLE=ON`. The short-prompt baseline is v0.1.39 built here (it also carries #861, which does nothing unless a request opts in); the other tables compare against the official v0.1.39 exe. For short prompts I sent chat requests with a ~100-token system prompt (reused from the prompt cache) and a different excerpt of `src/program/generate.cpp` each time, 1 token out; 30 requests per size, median of the last 28: | prompt read | v0.1.39 | this PR | |---|---:|---:| | 158 tokens | 425 ms | 365 ms (-14%) | | 315 tokens | 527 ms | 452 ms (-14%) | | 452 tokens | 593 ms | 524 ms (-12%) | The "dequant" phase of a ~500-token chunk went from ~300 ms to ~214 ms. For longer prompts I ran the official v0.1.39 exe and this build in alternation (A B A B), 1 warm-up and 3 runs each. They come out the same, which is expected since a 2.6K or 14.7K prompt takes the streamed walk: | | v0.1.39 | this PR | |---|---:|---:| | 2.6K prompt read | ~3,170 tok/s | ~3,130-3,173 tok/s | | 14.7K prompt read | ~6,460-6,480 tok/s | ~6,475 tok/s | To check that the numbers do not change, I compared per-token logprobs (`STRATA_LOGPOS`) on teacher-forced text, official exe against this build, alternating, two runs each. All four runs of each text gave the same perplexity: | text | v0.1.39 | this PR | target mismatches | |---|---:|---:|---:| | English, 512-token chunks x 32 (staged walk) | 8.9294 | 8.9294 | 0 | | Portuguese, 512-token chunks x 32 (staged walk) | 1.606 | 1.606 | 0 | | English, 2048-token chunks x 8 (streamed walk, control) | 6.8067 | 6.8067 | 0 | My first version crashed with an illegal memory access. After a staged expert, the next flush still covered the staged expert's position, whose group pointer was empty. The second commit (squashed here) starts the group after it.
En el sitio
Enlaces a install, modelos, releases.