Pull requests / #443
#443 kernels: add opt-in small-window GR and one-warp MMVQ experiments
closed · @CC-David-CC · 0 comentários · No GitHub
BenchmarksNVIDIA / CUDAModels & quantsDocumentationWindows
Descrição
Updated to upstream 0.1.34 (`1678de3`). Fresh targeted builds and regressions passed; the full fleet/context performance numbers below remain measurements of the prior 0.1.33 source. [Rebase checks](https://github.com/CC-David-CC/Strata-a5500/blob/perf/fleet-mmvq/docs/benchmarks/2026-10-02-upstream-sync-experimental.md) This independent, opt-in experiment is based directly on upstream 0.1.34 (`1678de3`). Hardware enablement and non-MTP serving have separate PRs. ## Changes - `STRATA_GR_DOWN_MAX4=1`: bound per-thread storage for windows of at most four tokens while preserving block dimensions, tile width, accumulation order and upstream plain/split/staged selection. Larger windows retain the existing path. - `STRATA_MMVQ_WARP1=1`: one warp per dense MMVQ output row. Floating-point association changes; this path uses a declared numerical-tolerance gate. Both are **off by default** and tested individually. Weights, quantization, expert selection and MTP threshold are unchanged. ## Measured benefit GSQ-RCO IQ3_S, Q8 KV, MTP4 at threshold0.5, long code/prose: | GPU / input | Option | Code generation | Prose generation | Interpretation | |---|---|---:|---:|---| | RTX PRO 6000 Blackwell / 64K | GR max4 | 223.37 -> 228.15 tok/s (+2.14%) | 160.21 -> 164.07 (+2.41%) | Exact output tokens matched | | RX 5500 XT 8GB / 8K | One-warp MMVQ | 16.24 -> 16.95 tok/s (+4.35%) | 13.08 -> 13.52 (+3.41%) | Output tokens and lengths differed | RTX PRO controls were untouched upstream A/B, bracketing the focused candidate. Three additional alternating default/max4 pairs reproduced the gain: code 1.77-2.49%, prose2.14-2.72%, with identical token IDs. Total-time improvement was only 0.89%/1.08% in the original comparison because prefill was unchanged. The 5500 XT used one default/warp1/default sequence. Default A/B tokens matched; the optimized outputs differed. Prose generated 1601 rather than 1453 tokens and took longer overall (218.57 versus 210.96 s). This is not an identical-output completion-time or quality claim. Its combined source tree is documented. Other single 8K + 512-counting-token GR pairs were: 4090 +0.30%, 3070 -0.03%, P4 +1.74%, and mapped-expert 7900 XTX -7.48%. Those single pairs do not establish small gains; the 7900 XTX result argues for retaining its default path. ## Correctness and limits - GR bitwise parity passed T1..8 across all three paths and write/injection combinations on all six GPUs. - Default MMVQ passed 182 real-tensor/column cases bitwise on all six GPUs, including the native output projection. - 5500 XT warp1 passed all 182 cases with maximum relative L1 **1.29335e-7** against the declared `1e-5` limit. Exact generated-token equality is not promised. - The final output-head coverage expansion is test-only commit `ac28acb`; measured engine binaries were unchanged. Source/binary identities are recorded. - Generated TTLCache classes passed bounded functional checks. These are not general model-quality or correctness guarantees for every input. [Scope, all results and reproduction](https://github.com/CC-David-CC/Strata-a5500/blob/perf/fleet-mmvq/docs/EXPERIMENTAL_KERNELS.md) | [Full matrix](https://github.com/CC-David-CC/Strata-a5500/blob/perf/fleet-mmvq/docs/benchmarks/2026-10-01-experimental-kernels.md) | [Machine-readable evidence](https://github.com/CC-David-CC/Strata-a5500/blob/perf/fleet-mmvq/docs/benchmarks/2026-10-01-experimental-kernels.json) Credit: [Niko1221/Strata](https://github.com/Niko1221/Strata) supplies the engine, expert cache and MTP. This patch changes only the two named launch/storage paths.
No site
Links install, modelos, releases.