Pull requests / #372
#372 Prompt path: an MMQ group's experts gathered in one launch, after one wait, released by one event (-9% on a 32K prompt, same output)
closed · @sergqwer · 0 comments · View on GitHub
Setup & installNVIDIA / CUDAModels & quantsWindows
Description
Every prompt chunk on a native pack runs the streamed walk: for each expert the compute stream waits for its copy, gathers it into the MMQ group buffer and records an event that gives its ring slot back. The gather takes ~3 us. Under WDDM, each wait and record also leaves ~10 us of GPU idle, even when the copy landed long before. An nsys trace of a 32K prompt shows: - 226 ms of idle between 22,659 back-to-back gathers; - the gather kernels themselves take 71 ms; - the host is 3-11 ms ahead of the GPU the whole time, so the idle is not the host being late. What changes: - **One launch per group:** `gather_native_group` / `copy16_group_kernel` copy a group's 16 experts in one launch, the same bytes to the same slots. - **One wait per group:** on the group's last streamed copy. The copy stream is in order, so it covers the others. - **One event per group:** it releases all of the group's ring slots. The issuer waits on `used[used_of[slot]]`, where `used_of` maps each slot to the event that released it. A later record of that event only waits longer. - **Skipped ring entries:** an entry the routing skipped inside an open group first gathers what the group holds so far. So no more than a group's entries are ever held back from the issuer, and a 96-slot ring cannot stall. - **Unchanged:** the products, their order, and the MMQ tail memsets. The staged path (chunks below 1024) and the FP16 path are as before. - **Opt-out:** `STRATA_PREFILL_GROUP_GATHER=0` gathers, waits and releases one expert at a time. ## Same output First-token logits are bit-identical with the switch on and off at 2K, 8K, 16K and 32K, and with a 96-slot ring (`STRATA_PREFILL_RING=96`) at 8K: - They were dumped with #276's `STRATA_DUMP_FIRST_LOGITS` in a test build. - The runs used a fixed `--expert-cache 14900`. With `auto` the slot count moves by ~100 from run to run with the desktop's free VRAM, and that alone changes a first token now and then, on main too. ## Speed Setup: IQ2_XS (ISTA), RTX 5090, 9950X3D, 128 GB, Windows 11, `--max-context 262144 --kv int8 --prefill auto --vram-reserve-mib 1500`, cache auto. The runs alternated between the switch off and on. | prompt | 0.1.31 | switch off | this PR | | --- | ---: | ---: | ---: | | 32K | 6,184 ms | 6,106 / 6,098 | **5,563 / 5,598** (-8.6%) | | 8K | 1,741 ms | 1,822 / 1,750 | **1,596 / 1,596** (-10%) | | 2K (fixed cache) | | 872 | **762** (-12.6%) | | 8K, 96-slot ring (fixed cache) | | 1,718 | **1,561** | On an NVFP4 pack (our fork, larger experts): 32K -5.2%, 16K -7.9%, 8K -11.1%. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.