Pull requests / #372

#372 Prompt path: an MMQ group's experts gathered in one launch, after one wait, released by one event (-9% on a 32K prompt, same output)

closed · @sergqwer · 0 评论 · 在 GitHub 查看

Setup & installNVIDIA / CUDAModels & quantsWindows

描述

Every prompt chunk on a native pack runs the streamed walk: for each expert the compute stream waits for its copy, gathers it into the MMQ group buffer and records an event that gives its ring slot back. The gather takes ~3 us. Under WDDM, each wait and record also leaves ~10 us of GPU idle, even when the copy landed long before.

An nsys trace of a 32K prompt shows:
- 226 ms of idle between 22,659 back-to-back gathers;
- the gather kernels themselves take 71 ms;
- the host is 3-11 ms ahead of the GPU the whole time, so the idle is not the host being late.

What changes:
- **One launch per group:** `gather_native_group` / `copy16_group_kernel` copy a group's 16 experts in one launch, the same bytes to the same slots.
- **One wait per group:** on the group's last streamed copy. The copy stream is in order, so it covers the others.
- **One event per group:** it releases all of the group's ring slots. The issuer waits on `used[used_of[slot]]`, where `used_of` maps each slot to the event that released it. A later record of that event only waits longer.
- **Skipped ring entries:** an entry the routing skipped inside an open group first gathers what the group holds so far. So no more than a group's entries are ever held back from the issuer, and a 96-slot ring cannot stall.
- **Unchanged:** the products, their order, and the MMQ tail memsets. The staged path (chunks below 1024) and the FP16 path are as before.
- **Opt-out:** `STRATA_PREFILL_GROUP_GATHER=0` gathers, waits and releases one expert at a time.

## Same output

First-token logits are bit-identical with the switch on and off at 2K, 8K, 16K and 32K, and with a 96-slot ring (`STRATA_PREFILL_RING=96`) at 8K:
- They were dumped with #276's `STRATA_DUMP_FIRST_LOGITS` in a test build.
- The runs used a fixed `--expert-cache 14900`. With `auto` the slot count moves by ~100 from run to run with the desktop's free VRAM, and that alone changes a first token now and then, on main too.

## Speed

Setup: IQ2_XS (ISTA), RTX 5090, 9950X3D, 128 GB, Windows 11, `--max-context 262144 --kv int8 --prefill auto --vram-reserve-mib 1500`, cache auto. The runs alternated between the switch off and on.

| prompt | 0.1.31 | switch off | this PR |
| --- | ---: | ---: | ---: |
| 32K | 6,184 ms | 6,106 / 6,098 | **5,563 / 5,598** (-8.6%) |
| 8K | 1,741 ms | 1,822 / 1,750 | **1,596 / 1,596** (-10%) |
| 2K (fixed cache) | | 872 | **762** (-12.6%) |
| 8K, 96-slot ring (fixed cache) | | 1,718 | **1,561** |

On an NVFP4 pack (our fork, larger experts): 32K -5.2%, 16K -7.9%, 8K -11.1%.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。