Pull requests / #443

#443 kernels: add opt-in small-window GR and one-warp MMVQ experiments

closed · @CC-David-CC · 0 comentarios · En GitHub

BenchmarksNVIDIA / CUDAModels & quantsDocumentationWindows

Descripción

Updated to upstream 0.1.34 (`1678de3`). Fresh targeted builds and regressions passed; the full fleet/context performance numbers below remain measurements of the prior 0.1.33 source.

[Rebase checks](https://github.com/CC-David-CC/Strata-a5500/blob/perf/fleet-mmvq/docs/benchmarks/2026-10-02-upstream-sync-experimental.md)

This independent, opt-in experiment is based directly on upstream
0.1.34 (`1678de3`). Hardware enablement and non-MTP serving have separate PRs.

## Changes

- `STRATA_GR_DOWN_MAX4=1`: bound per-thread storage for windows of at most four
  tokens while preserving block dimensions, tile width, accumulation order and
  upstream plain/split/staged selection. Larger windows retain the existing path.
- `STRATA_MMVQ_WARP1=1`: one warp per dense MMVQ output row. Floating-point
  association changes; this path uses a declared numerical-tolerance gate.

Both are **off by default** and tested individually. Weights, quantization,
expert selection and MTP threshold are unchanged.

## Measured benefit

GSQ-RCO IQ3_S, Q8 KV, MTP4 at threshold0.5, long code/prose:

| GPU / input | Option | Code generation | Prose generation | Interpretation |
|---|---|---:|---:|---|
| RTX PRO 6000 Blackwell / 64K | GR max4 | 223.37 -> 228.15 tok/s (+2.14%) | 160.21 -> 164.07 (+2.41%) | Exact output tokens matched |
| RX 5500 XT 8GB / 8K | One-warp MMVQ | 16.24 -> 16.95 tok/s (+4.35%) | 13.08 -> 13.52 (+3.41%) | Output tokens and lengths differed |

RTX PRO controls were untouched upstream A/B, bracketing the focused candidate.
Three additional alternating default/max4 pairs reproduced the gain: code
1.77-2.49%, prose2.14-2.72%, with identical token IDs. Total-time improvement was
only 0.89%/1.08% in the original comparison because prefill was unchanged.

The 5500 XT used one default/warp1/default sequence. Default A/B tokens matched;
the optimized outputs differed. Prose generated 1601 rather than 1453 tokens
and took longer overall (218.57 versus 210.96 s). This is not an identical-output
completion-time or quality claim. Its combined source tree is documented.

Other single 8K + 512-counting-token GR pairs were: 4090 +0.30%, 3070 -0.03%,
P4 +1.74%, and mapped-expert 7900 XTX -7.48%. Those single pairs do not establish
small gains; the 7900 XTX result argues for retaining its default path.

## Correctness and limits

- GR bitwise parity passed T1..8 across all three paths and write/injection
  combinations on all six GPUs.
- Default MMVQ passed 182 real-tensor/column cases bitwise on all six GPUs,
  including the native output projection.
- 5500 XT warp1 passed all 182 cases with maximum relative L1 **1.29335e-7**
  against the declared `1e-5` limit. Exact generated-token equality is not promised.
- The final output-head coverage expansion is test-only commit `ac28acb`;
  measured engine binaries were unchanged. Source/binary identities are recorded.
- Generated TTLCache classes passed bounded functional checks. These are not
  general model-quality or correctness guarantees for every input.

[Scope, all results and reproduction](https://github.com/CC-David-CC/Strata-a5500/blob/perf/fleet-mmvq/docs/EXPERIMENTAL_KERNELS.md)
| [Full matrix](https://github.com/CC-David-CC/Strata-a5500/blob/perf/fleet-mmvq/docs/benchmarks/2026-10-01-experimental-kernels.md)
| [Machine-readable evidence](https://github.com/CC-David-CC/Strata-a5500/blob/perf/fleet-mmvq/docs/benchmarks/2026-10-01-experimental-kernels.json)

Credit: [Niko1221/Strata](https://github.com/Niko1221/Strata) supplies the engine,
expert cache and MTP. This patch changes only the two named launch/storage paths.

En el sitio

Enlaces a install, modelos, releases.