Pull requests / #1033

#1033 prefill: gather GPU-resident experts in groups on short prompts too

closed · @brenoperucchi · 0 comentários · No GitHub

BenchmarksSetup & installNVIDIA / CUDAModels & quantsWindows

Descrição

#372 (in 0.1.38) gathers an MMQ group's experts in one launch, but only in the streamed walk. A chunk below
`stream_all_min()` (1024 tokens) still takes the staged walk, where every expert gets its own `gather_native` launch.
@BlueKingMuch noticed this in #372: the 923-token tail of a 64K prompt was not affected.

On our box most short requests are 150-600 tokens with almost every routed expert already in the GPU cache, so this
per-expert launch is where their time goes. With `STRATA_PREFILL_TIMING=1`, a ~450-token prompt spends 50-55% of its
GPU timeline in the "dequant" phase, which on the MMQ path is the gather.

The staged walk was left out of #372 because a staged expert sits in the 8-slot staging ring, which is smaller than a
group of 16, so a group cannot hold its slots. Resident experts have no ring slot, no wait and no release. This patch
groups only those:

- a resident expert joins the open group (`gather_native_group`, the same bytes into the same group slots);
- a staged expert first flushes what the group holds, then is gathered alone as before, and the group restarts after
  it;
- products, their order, the MMQ tail memsets and the streamed walk are unchanged;
- `STRATA_PREFILL_GROUP_GATHER=0` still turns grouping off everywhere.

One file, `src/prefill/prefill.cpp`, +17/-2.

Tested on an RTX 5090 32 GB, Ryzen 9 5950X, 96 GB DDR4, Windows 11 (WDDM), CUDA 13.0, with the Swift IQ3_XXS native pack,
`--expert-cache auto`, `--prefill auto:32768`, `--kv int8`, `--max-context 32768`. My builds use the
setup.py flags plus `STRATA_PORTABLE=ON`. The short-prompt baseline is v0.1.39 built here (it also carries #861, which
does nothing unless a request opts in); the other tables compare against the official v0.1.39 exe.

For short prompts I sent chat requests with a ~100-token system prompt (reused from the prompt cache) and a different excerpt of
`src/program/generate.cpp` each time, 1 token out; 30 requests per size, median of the last 28:

| prompt read | v0.1.39 | this PR |
|---|---:|---:|
| 158 tokens | 425 ms | 365 ms (-14%) |
| 315 tokens | 527 ms | 452 ms (-14%) |
| 452 tokens | 593 ms | 524 ms (-12%) |

The "dequant" phase of a ~500-token chunk went from ~300 ms to ~214 ms.

For longer prompts I ran the official v0.1.39 exe and this build in alternation (A B A B), 1 warm-up and 3 runs each.
They come out the same, which is expected since a 2.6K or 14.7K prompt takes the streamed walk:

| | v0.1.39 | this PR |
|---|---:|---:|
| 2.6K prompt read | ~3,170 tok/s | ~3,130-3,173 tok/s |
| 14.7K prompt read | ~6,460-6,480 tok/s | ~6,475 tok/s |

To check that the numbers do not change, I compared per-token logprobs (`STRATA_LOGPOS`) on teacher-forced text,
official exe against this build, alternating, two runs each. All four runs of each text gave the same perplexity:

| text | v0.1.39 | this PR | target mismatches |
|---|---:|---:|---:|
| English, 512-token chunks x 32 (staged walk) | 8.9294 | 8.9294 | 0 |
| Portuguese, 512-token chunks x 32 (staged walk) | 1.606 | 1.606 | 0 |
| English, 2048-token chunks x 8 (streamed walk, control) | 6.8067 | 6.8067 | 0 |

My first version crashed with an illegal memory access. After a staged expert, the next flush still covered the
staged expert's position, whose group pointer was empty. The second commit (squashed here) starts the group after it.

No site

Links install, modelos, releases.