Pull requests / #1033

#1033 prefill: gather GPU-resident experts in groups on short prompts too

closed · @brenoperucchi · 0 comentarios · En GitHub

BenchmarksSetup & installNVIDIA / CUDAModels & quantsWindows

Descripción

#372 (in 0.1.38) gathers an MMQ group's experts in one launch, but only in the streamed walk. A chunk below
`stream_all_min()` (1024 tokens) still takes the staged walk, where every expert gets its own `gather_native` launch.
@BlueKingMuch noticed this in #372: the 923-token tail of a 64K prompt was not affected.

On our box most short requests are 150-600 tokens with almost every routed expert already in the GPU cache, so this
per-expert launch is where their time goes. With `STRATA_PREFILL_TIMING=1`, a ~450-token prompt spends 50-55% of its
GPU timeline in the "dequant" phase, which on the MMQ path is the gather.

The staged walk was left out of #372 because a staged expert sits in the 8-slot staging ring, which is smaller than a
group of 16, so a group cannot hold its slots. Resident experts have no ring slot, no wait and no release. This patch
groups only those:

- a resident expert joins the open group (`gather_native_group`, the same bytes into the same group slots);
- a staged expert first flushes what the group holds, then is gathered alone as before, and the group restarts after
  it;
- products, their order, the MMQ tail memsets and the streamed walk are unchanged;
- `STRATA_PREFILL_GROUP_GATHER=0` still turns grouping off everywhere.

One file, `src/prefill/prefill.cpp`, +17/-2.

Tested on an RTX 5090 32 GB, Ryzen 9 5950X, 96 GB DDR4, Windows 11 (WDDM), CUDA 13.0, with the Swift IQ3_XXS native pack,
`--expert-cache auto`, `--prefill auto:32768`, `--kv int8`, `--max-context 32768`. My builds use the
setup.py flags plus `STRATA_PORTABLE=ON`. The short-prompt baseline is v0.1.39 built here (it also carries #861, which
does nothing unless a request opts in); the other tables compare against the official v0.1.39 exe.

For short prompts I sent chat requests with a ~100-token system prompt (reused from the prompt cache) and a different excerpt of
`src/program/generate.cpp` each time, 1 token out; 30 requests per size, median of the last 28:

| prompt read | v0.1.39 | this PR |
|---|---:|---:|
| 158 tokens | 425 ms | 365 ms (-14%) |
| 315 tokens | 527 ms | 452 ms (-14%) |
| 452 tokens | 593 ms | 524 ms (-12%) |

The "dequant" phase of a ~500-token chunk went from ~300 ms to ~214 ms.

For longer prompts I ran the official v0.1.39 exe and this build in alternation (A B A B), 1 warm-up and 3 runs each.
They come out the same, which is expected since a 2.6K or 14.7K prompt takes the streamed walk:

| | v0.1.39 | this PR |
|---|---:|---:|
| 2.6K prompt read | ~3,170 tok/s | ~3,130-3,173 tok/s |
| 14.7K prompt read | ~6,460-6,480 tok/s | ~6,475 tok/s |

To check that the numbers do not change, I compared per-token logprobs (`STRATA_LOGPOS`) on teacher-forced text,
official exe against this build, alternating, two runs each. All four runs of each text gave the same perplexity:

| text | v0.1.39 | this PR | target mismatches |
|---|---:|---:|---:|
| English, 512-token chunks x 32 (staged walk) | 8.9294 | 8.9294 | 0 |
| Portuguese, 512-token chunks x 32 (staged walk) | 1.606 | 1.606 | 0 |
| English, 2048-token chunks x 8 (streamed walk, control) | 6.8067 | 6.8067 | 0 |

My first version crashed with an illegal memory access. After a staged expert, the next flush still covered the
staged expert's position, whose group pointer was empty. The second commit (squashed here) starts the group after it.

En el sitio

Enlaces a install, modelos, releases.