Pull requests / #363

#363 experts: smaller and fewer launches for the verify window's PCIe call

closed · @BlueKingMuch · 0 コメント · GitHub で見る

BenchmarksNVIDIA / CUDAModels & quantsWindows

本文

Every verify window calls `native_expert_grouped` twice per layer: for the experts in VRAM and for the ones that come over PCIe. Both calls launched gate/up and down with one block row per *possible* group (`cap_groups` = window tokens x top-k) and ran SwiGLU and the q8_1 quantization as two kernels over all `cap_entries`. The PCIe call has no group at all at `pcie_frac` 0 and rarely more than one otherwise, yet every layer paid `cap_groups` x (2 `n_ff` / 8 + `n_embd` / 8) blocks for it that read `n_groups` and returned (about 17k with this model's 3.4-token windows), plus two passes over entries that were not its own. The graph is captured once per session and a request may still set `pcie_frac` afterwards, so the call stays in the graph (the graph without the PCIe path was left out when #109 was merged); this PR makes it cheap instead.

## What changes

- **Group stride:** the four grouped kernels (`native_gu_kernel`, `native_down_kernel` and #242's `_multi` variants) run group `g` in block row `g % gridDim.y` (rows `y, y + gridDim.y, ...` below `*n_groups`). The verify window's PCIe call launches `kPcieGroupRows` = 4 block rows; the VRAM call and every other caller keep one per possible group (`grid_groups` defaults to 0 = `cap_groups`). A group's rows are computed by the same warp code in the same order whichever block row runs it, so the results do not depend on the grid.
- **SwiGLU and q8_1 in one kernel**, over the call's own entries `[grp_start[0], grp_start[n_groups])`. The SwiGLU product is rounded on its own before the q8_1 block sums, as it was when it went through memory (`__fmul_rn`). The quantizer body is `quantize_q8_1_kernel`'s, factored out into a device function; that kernel's PTX is unchanged.
- `STRATA_GROUPED_V1=1` keeps the old launches, for A/B timing.

It is on by default since it changes no bit. If you prefer it opt-in, that is the default of `g_grouped_v1`.

## Exactness

- `native_grouped_parity` (a GPU ctest next to `iq_multi_parity`) runs the new launches against `STRATA_GROUPED_V1`'s and compares the outputs bit for bit: every gate/up format with every down format the dispatcher takes (gate/up IQ2_XXS, IQ2_XS, IQ3_XXS, IQ3_S, IQ2_S, IQ4_XS, IQ1_M, Q2_0, Q4_K, Q5_K, Q8_0; down IQ4_NL, IQ4_XS, Q2_0, Q5_1, Q8_0), calls of 0 to `cap` groups, entries from 0 and from past 0 (the PCIe call's entries follow the VRAM call's), scattered destinations, stale scratch, and `grid_groups` 0, 1, 2, 3, 4 and above `cap`. The rows a call must not write are compared too (both runs start from the same NaN pattern). On this branch (0.1.31): 0 failures.
- End to end: requests without anything timing-dependent (no prompt-lookup drafts, no expert swaps, `pcie_frac=0`) give the same 256 tokens in the same number of rounds with and without `STRATA_GROUPED_V1=1`, greedy and sampled (T 1.0, top_p 0.95, top_k 20, seed 1).

## Measured

RTX 4080 SUPER 32 GB (sm_89), Ryzen 7 5800X3D, 64 GB DDR4, PCIe 3.0 x16, Windows 11, CUDA 13.3.

**The calls alone**, on this branch: `native_grouped_parity --bench` captures a window's 48 calls in a graph, each layer with its own experts (2560 x 768, IQ3_S gate/up, IQ4_XS down), so the weights stream from DRAM as in the engine. The two variants' launches alternate; medians of 40 launches each:

| call | before (`STRATA_GROUPED_V1=1`) | this PR |
|---|---|---|
| PCIe call, no group (`pcie_frac` 0) | 16.51 us | 4.01 us |
| PCIe call, one group of two entries | 19.11 us | 10.71 us |
| VRAM call, 28 groups / 40 entries | 149.72 us | 148.95 us |

Kernel by kernel (Nsight Systems on the same bench), the VRAM call's gate/up and down take the same time in both variants (86.2 / 85.8 us and 54.7 / 54.8 us), and the fused SwiGLU + q8_1 kernel takes 2.06 us against 1.41 + 1.47 us for the two it replaces.

One thing to watch when timing this on a desktop card: under sustained load this one steps its clock down after a fraction of a second (gate/up went from 80 to 89 us within one run), so timing one variant after the other can show a difference of that size that has nothing to do with the change. That is why the bench alternates them.

**In the engine:** 0.1.31 (9259cad) with this commit only, plus two measurement aids that change no computed value: #320 (so the profile's columns add up) and a local patch that lets a request ask for a fixed length and for no timing-dependent parts (`ignore_eos`, `suffix=0`, `adapt=0`). IQ3_S, `--max-context 262144`, int8 KV, 11055 experts in VRAM; every comparison within that build (`STRATA_GROUPED_V1` unset against `=1`).

GPU stages (`STRATA_VERIFY_PROFILE=1`, 8K context, ms per verify window, DeltaNet and attention layers summed):

| | before (`STRATA_GROUPED_V1=1`) | this PR |
|---|---|---|
| experts over PCIe | 0.67 | 0.20 |
| experts in VRAM | 4.56 | 4.52 |
| waiting for the plan and the CPU experts | 1.45 | 1.88 |
| whole window | 19.99 | 19.91 |

The PCIe call loses its 0.47 ms in every run I made, also in a 0.1.30-based build and next to other open PRs. How much of it the window keeps depends on what the GPU waits for next: in these runs it mostly waited longer for the host side instead (+0.43 ms for the plan and the CPU experts), so the window got 0.08 ms shorter; in the 0.1.30-based run it kept all of it (18.97 -> 18.50 ms).

The deterministic requests above (the same 256 tokens in 97 / 103 rounds; with expert swaps off the CPU experts take about 7.5 ms per round): round 24.79 -> 24.49 ms greedy and 25.13 -> 24.85 ms sampled (-0.30 / -0.28 ms, +1.1 % tokens/s), GPU wait per round -0.31 / -0.26 ms.

Decode, 5 requests of 256 tokens per cell, prompt-lookup drafts and expert swaps on (8K greedy / sampled, 64K greedy / sampled):

| | before | this PR |
|---|---|---|
| GPU wait per round, ms | 14.76 / 14.86 / 16.35 / 16.43 | 14.60 / 14.77 / 15.89 / 16.10 |
| tokens/s | 115.5 / 114.3 / 111.9 / 114.1 | 118.0 / 113.9 / 119.6 / 116.4 |
| `--pcie-frac 0.28`: GPU wait per round, ms | 14.85 / 14.94 / 16.46 / 16.50 | 14.41 / 14.57 / 16.05 / 16.05 |
| `--pcie-frac 0.28`: tokens/s | 115.5 / 109.8 / 115.7 / 114.5 | 113.3 / 112.7 / 117.9 / 112.6 |

The GPU wait per round drops in all eight cells (by 0.09 to 0.46 ms). The tokens/s stay within the scatter of single requests (104 to 130 within one cell here; without the SSD keep-alive of #317 some requests also waited for SSD reads): +2.2 / -0.3 / +6.9 / +2.0 % at `pcie_frac` 0, -1.9 / +2.6 / +1.9 / -1.7 % at 0.28.

In short: a fixed 0.47 ms of GPU time per window, tokens and bits unchanged. What it buys in tokens/s depends on whether the GPU or the CPU experts set the pace on a given machine; about 1 % in the deterministic runs here.

To reproduce: `STRATA_GROUPED_V1=1` against unset, `STRATA_VERIFY_PROFILE=1` for the stages, `native_grouped_parity --bench` for the calls alone.

関連リンク

インストール・モデル・リリースへの站内リンク。