Pull requests / #1181

#1181 perf(expert): the verify window's GPU plan in O(n) instead of O(n^2)

open · @ZhongUncle · 0 commentaires · Sur GitHub

BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsWindows

Description


The verify window's GPU plan is built with O(n²) scans, and the GPU spin-waits on flag A while it is built: the plan is ["decided and published FIRST so the GPU starts while the CPU works"](https://github.com/Niko1221/Strata/blob/82f46a8c8f475f001ad76d92f58f4a4f8ffb0253/src/core/expert_source.cpp#L2445), and flag A rises only after it, at [`P.publish`](https://github.com/Niko1221/Strata/blob/82f46a8c8f475f001ad76d92f58f4a4f8ffb0253/src/core/expert_source.cpp#L2540). Plan time is therefore GPU idle time on the flag-A critical path, even where its share of the round is too small to show end-to-end (see Limits). The quadratic parts: a pairwise first-occurrence scan ([`expert_source.cpp:2462`](https://github.com/Niko1221/Strata/blob/82f46a8c8f475f001ad76d92f58f4a4f8ffb0253/src/core/expert_source.cpp#L2462)), three full re-scans of the window per distinct expert (kind assignment, VRAM and PCIe group entries), and a `std::find` dedup of the CPU-miss prefetch ([`expert_source.cpp:2598`](https://github.com/Niko1221/Strata/blob/82f46a8c8f475f001ad76d92f58f4a4f8ffb0253/src/core/expert_source.cpp#L2598)).

This PR rewrites all of it as three linear passes with output-identical results. It deliberately does **not** change what the plan decides — the tiers, the PCIe share and the group order are exactly the old ones. The case for taking it is structural: no behavior change (pinned by a 4,020-case differential test, no GPU needed), and the quadratic term it removes grows with the window, so a bigger `--spec` widens the saving.

**TL;DR**

- ~3.1–45.7K branchy iterations per window-layer (3×10 to 12×10 windows, by n(n−1)/2 + 3·distinct·n) collapse to three linear passes
- standalone planner: 2.3–3.2× for the windows seen in the engine A/B (mean ≈3×10, max 6×10), 4.8× near the 128-entry cap
- engine `--stats`: the dispatch `plan` figure halves in both rig states seen here (four interleaved pairs); decode unchanged, as expected — the ≈0.3–0.4% gain is below the ±0.7% run-to-run spread
- output-identical, proven without a GPU (4,020/4,020; the reference is diff-verified against the old code) — details in [Findings](#findings), caveats in [Limits](#limits)

### What changes

The planner is extracted as `detail::window_gpu_plan` (the header's existing `detail` test-seam pattern) and rewritten:

- one pass groups each entry with its expert's first occurrence — open addressing over the routing ids, 256 slots — replacing the pairwise `first_of` scan
- one pass over the distinct experts decides the tiers — the peer/helper-GPU questions are answered once per distinct expert instead of once per entry and again per group
- one pass emits the groups with per-group cursors, replacing the three per-distinct-expert re-scans
- the CPU-miss prefetch dedup uses the same open-addressed set instead of `std::find`

Bounds are the old ones: the call site already guards `n <= kMaxWindowEntries` (128) and falls back to the simple no-plan loop beyond it, so the 256-slot table never exceeds load ½; out-of-range ids group by the same equality the pairwise scan used and stay kind −1, refused later by the pool loop with the same message. Unchanged by construction: group order (routing order of first occurrence), entry order within a group, counts, kinds.

### Measured results

Measured on:

- RTX 3080 20 GB (driver 595.91.07, CUDA 13.2)
- Xeon E5-2680 v4 (no AVX512)
- 62 GB RAM
- Ubuntu 24.04
- Qwen3.8-Flash-Next IQ2_XS pack (48 layers × 512 experts = 24,576 possible pairs)
- `--spec 4 --spec-min-p 0.5` + the MTP drafter (verify window up to 6 tokens)

**Standalone planner, ns per window-layer (512 experts, ~70% resident, 1,000 pre-generated windows cycled through, 200,000 rounds per arm):**

| window | old | new | speedup (×) |
|---|---|---|---|
| 3×10 (30 entries, 29 avg distinct) | 1796 | 790 | 2.3 |
| 4×10 (40 entries, 39 avg distinct) | 2686 | 1043 | 2.6 |
| 5×10 (50 entries, 48 avg distinct) | 3854 | 1311 | 2.9 |
| 6×10 (60 entries, 57 avg distinct) | 5103 | 1576 | 3.2 |
| 8×10 (80 entries, 74 avg distinct) | 8078 | 2081 | 3.9 |
| 12×10 (120 entries, 107 avg distinct) | 15283 | 3204 | **4.8** |

(Per 48-layer window that is 86 → 38 µs at 3×10 up to 734 → 154 µs at 12×10 of GPU wait. The window grows with `--spec`: 6×10 is this config's maximum, 12×10 is near the 128-entry cap. The gap widens with window size because the old cost grows with distinct × n.)

**Engine A/B — `dispatch` line sections in ms per window-round, mean of two replicated pairs per rig state (four interleaved pairs total; raw lines in the appendix):**

| section | baseline (slow) | patched (slow) | baseline (fast) | patched (fast) |
|---|---|---|---|---|
| plan | 0.19 | **0.09** | 0.14 | **0.07** |
| activation quantize | 0.68 | 0.68 | 0.42 | 0.42 |
| jobs | 0.02 | 0.02 | 0.01 | 0.01 |
| run | 7.56 | 7.66 | 4.87 | 4.92 |

(decode: baseline 95.7–97.1, patched 95.8–96.8 tok/s over four runs each — ranges overlap; durations in the appendix.)

### Findings

1. **The new planner's output is identical to the old one's.** `tests/core/expert_window_plan_test.cpp` drives `detail::window_gpu_plan` and a reference copy of the O(n^2) planner with the same inputs and compares every positional field — kinds, group `ptr`/`start`/`start2`/`dst`/`tok`, the group counts, `nmiss`, the DMA source list — on 20 targeted corners (the PCIe share's last-m boundary, the staging cap, peer/helper-held experts, out-of-range ids, `pcie_mode` 0/1/2, per-slot cache offsets) and 4,000 randomized windows: identical in all 4,020. The reference is the replaced block with only the mechanical adaptation to the injected predicates, diff-verified against the base commit; the production predicates are the old inline expressions verbatim.
2. **Standalone plan time drops 3.2× at this config's max window, 4.8× near the entry cap.** 5,103 → 1,576 ns per window-layer at 6×10 (245 → 76 µs of GPU wait per 48-layer window); the old cost grows with distinct × n, the new one with n — measured across six window sizes from 2.3× at 3×10 to 4.8× at 12×10 (table above; the in-engine figure is finding 3). (6 tokens is the cap at `--spec 4` because the default suffix drafter may lengthen a window by 2 — [`generate.cpp:2091`](https://github.com/Niko1221/Strata/blob/82f46a8c8f475f001ad76d92f58f4a4f8ffb0253/src/program/generate.cpp#L2091); the log's "rounds of 6" in the appendix confirms it is reached.)
3. **In the engine, the dispatch `plan` figure halves and nothing else moves.** 0.19 → 0.09 and 0.14 → 0.07 ms per window-round, replicated in both states; activation-quantize, jobs and run are unchanged within noise state-for-state (the run section's pair deltas are −0.01 to +0.22 ms).

   The engine's absolute numbers sit below the micro-benchmark's 48-layer extrapolation for two reasons. The `plan` segment times more than the planner — `begin_layer`, usage accounting, the flag-raising `P.publish` and the PCIe `P.fetch` issue share the segment and are untouched, so the ~0.07 ms that remains is those parts plus the now-small planner. And the A/B windows are smaller than the benchmark's headline arms: the logged lookup counts put the mean window at ≈3.3×10 slow-state and ≈2.4×10 fast-state (308,226 lookups ÷ 195 rounds ÷ 48 layers ≈ 33 and 283,774 ÷ 244 ÷ 48 ≈ 24 entries per window-layer — PCIe-served misses excluded; the draft counts agree: 1 + 458/195 ≈ 3.3 and 1 + 355/244 ≈ 2.4 tokens). The per-layer saving (105 µs/round ÷ 48 ≈ 2.2 µs slow-state, 67 ÷ 48 ≈ 1.4 fast-state) is 1.4–2.2× the benchmark's ≈1.0 µs at 3×10 — the expected direction, since the benchmark runs warm (L3-resident working set, regular branches, fixed residency — see Limits) while the engine is cold and irregular, and plan cost is convex in window size, so a spread of real windows costs more than the mean window. The micro-benchmark is the scaling story; the engine figure is the realistic one.
4. **Decode tok/s is unchanged, and that is the expected outcome, not hidden work.** Within each pair the two arms produce identical progress lines (greedy sampling; every logged counter matches — same rounds, same acceptance, same expert lookups, see the appendix): 195 rounds in the slow state, 244 in the fast, so a round is 27.3 ms slow-state and 21.8 fast-state. The plan saving (105 µs slow-state, 67 fast-state) is an expected ≈0.4% / ≈0.3% decode gain, below the ±0.7% run-to-run spread of the four runs (95.7–97.1 tok/s). The removed plan time is real dead time on the flag-A critical path; it becomes visible end-to-end where its share is larger (bigger `--spec` windows, slower CPUs — see Limits).

### Validation

- **What was compared:** the extracted planner against a reference copy of the O(n^2) code, same inputs, every positional field (finding 1); engine `--stats` dispatch lines on four interleaved baseline/patched pairs, 512 generated tokens after an 873-token prompt (finding 3)
- **Reference fidelity:** diff-verified against the base commit (82f46a8c) — see finding 1
- **Result:** 4,020/4,020 identical; `ctest -R expert_window_plan` passes in ~0.6 s with no GPU
- **Full suite:** `ctest` green except `ple_parity` and `expert_multi_test`, which fail identically on the base commit on this host (missing model files, no AVX512)

### Limits

- The planner is host code: on a faster CPU the absolute win shrinks, slower CPUs should see more — an expectation from the algorithm, **not** a measurement on a second rig. Bigger windows do see more, measured (2.3× at 3×10 → 4.8× at 12×10, table above)
- Decode tok/s does **not** measurably improve on this rig — the expected gain (≈0.3–0.4%) is below the run-to-run spread (finding 4)
- The micro-benchmark cycles 1,000 pre-generated windows per arm (no single window's branches can be learned), but its working set still fits this CPU's L3 and residency is a fixed ~70% — treat the ratios as the scaling story; the engine A/B is the realistic figure
- One model, one quant, one GPU
- The engine baseline binary was built from the pre-rewrite main (1735d647, still fetchable by SHA on GitHub); its build inputs are hash-identical to this PR's base — `src` 9f4e26be, `include` d4bf5e4b, `cmake` 478c0fb6, `CMakeLists.txt` cd605177 in both commits (the rewrite touched no C++ or build files)

### Reproduce

```bash
# equivalence, no GPU
ctest -R expert_window_plan

# engine A/B (host with the model; pretokenize via the pack's chat template)
./engine/strata --pack Strata-data/packs/iq2_xs --native <shard1.gguf> --ple-gguf <shard2.gguf> \
  --expert-profile data/expert-profile.bin --expert-cache auto --prefill auto \
  --spec 4 --spec-min-p 0.5 --mtp Strata-data/mtp/rt \
  --max-context 24576 --kv int8 --kv-resident 32768 \
  --tokens-file PROMPT.tokens --max-new 512 --stats
# the dispatch line's plan figure is what this PR reduces; interleave baseline and patched runs
```

The standalone micro-benchmark is **not** part of this PR — it is the test's driver with a timing loop; its parameters are in the appendix.

<details>
<summary>Raw `--stats` lines and the micro-benchmark</summary>

Engine A/B, all eight runs — 873-token task prompt, 512 generated tokens, interleaved baseline/patched pairs, slow/fast = the two rig states seen across the four pairs (clock and trajectory both differ between them; supports findings 3 and 4):

```
baseline r1: dispatch  plan 0.194  activation quantize 0.681  jobs 0.021  run 7.518 ms/round — decode 96.29 tok/s (5317.1 ms)
patched  r1: dispatch  plan 0.089  activation quantize 0.688  jobs 0.021  run 7.737 ms/round — decode 95.78 tok/s (5345.4 ms)
baseline r2: dispatch  plan 0.135  activation quantize 0.419  jobs 0.014  run 4.913 ms/round — decode 96.29 tok/s (5317.4 ms)
patched  r2: dispatch  plan 0.068  activation quantize 0.422  jobs 0.014  run 4.910 ms/round — decode 96.27 tok/s (5318.4 ms)
baseline r3: dispatch  plan 0.135  activation quantize 0.417  jobs 0.014  run 4.822 ms/round — decode 97.11 tok/s (5272.5 ms)
patched  r3: dispatch  plan 0.068  activation quantize 0.418  jobs 0

Sur le site

Liens install, modèles, releases.