Pull requests / #1742

#1742 CUDA SM86: opt-in one-superblock prefetch build

open · @InB4DevOps · 0 comentários · No GitHub

BenchmarksSetup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

Descrição

## Summary

Add the off-by-default **`STRATA_CUDA_SM86_PREFETCH_ONE` CMake option**. With an
explicit CUDA architecture 86 target, it builds a CUDA-only specialization for the
existing native fused path. Runtime opt-in remains `STRATA_PF_FUSED=1 STRATA_PF_PREFETCH_ONE=1`.
Gate/up weights are prefetched one superblock ahead instead of two.

**Updated scope:** the initial draft edited shared CUDA/HIP source. That edit and
the shared C++ test change have now been completely reverted. The new translation
unit is selected only inside the CUDA CMake branch, with an additional runtime
SM86 check. HIP's implementation/build branch and SYCL's build remain unchanged;
there is no HIP change to build. Other CUDA devices use the original implementation.
SM86 includes more cards than the 3060; measurements remain 3060-specific. Windows
and other GPU hardware have not been tested.

## What changed

- CUDA-only kernel specialization for gate/up, reusing conversion/routing helpers
  and the fallback by including the unchanged original implementation. Down keeps
  its original instance. The specialized body must retain the original arithmetic.
- A fresh-process numerical parity wrapper for six format pairs and both tile sizes.
- Scoped benchmark report, exact public-source prompts, per-pair timing/output-hash
  evidence, and final isolated validation results.

No activation-stage, tile/grid, residency, decode, model/setup-default or HIP-kernel
policy change. This branch is independently based on `fb58e0db` and contains none
of the other experimental engine patches.

## Measured evidence

RTX 3060 **12 GB / 100 W**, i7-12700KF AVX2, Linux, CUDA 13.4.92, native IQ3_XXS,
explicit 4096-token prefill and static residency. Six balanced pairs per realistic
prompt, 128 generated tokens:

| Prompt | Input tokens | Median paired prompt-speed gain | Faster pairs |
|---|---:|---:|---:|
| Parser/test review | 12,247 | +0.548% | 5/6 |
| Runtime documentation | 20,324 | +0.813% | 6/6 |
| C++ prefill review | 30,757 | +0.802% | 6/6 |

All 18 pairs had identical token IDs and draft counts, with within-arm repeatability.
Three independent fixed-work process pairs per shape gave **+1.71–6.52%** across
twelve format/size combinations; all 36 shape/process pairs were faster with matching hashes.

These are **prototype measurements of the same policy**, not inflated to represent
the separate MMQ-to-fused speedup. The isolated extraction's tests/smoke pairs are
reported separately. A/A controls had near-zero medians but individual outliers up
to about +/-1.5%; all outliers remain in the evidence. No broad GPU/model/default
or quality-equivalence claim is made.

## Validation

- CUDA SM86 Release: default-off build and opt-in build of `strata` and `prefill_fused_iq_test`.
- Configuration rejects the opt-in without architecture 86 and rejects combining the two build variants.
- Shared kernel/test files verified byte-identical to upstream; HIP CMake branch and SYCL build file unchanged.
- Isolated off/on outputs: six formats × both tiles, plus prior disabled-build
  default-build fingerprint comparison; existing FP64/MMQ numerical checks retained.
- Isolated two-pair whole-engine checks on the same three realistic prompts;
  token IDs, draft counts and repeatability checked.
- `git diff --check`.
- Non-SM86 dispatch predicate checked at compile time; other GPU hardware/Windows runtime untested.

Details and commands: `docs/benchmarks/2026-10-09-rtx3060-prefetch-one/README.md`.
Current validation: `cuda-only-evidence.json` and `cuda-only-parity.log`; historical
prototype/initial extraction: `evidence.json` and `isolated-parity.log`. Exact prompts are included.

The separate IQ3-only stage-buffering proposal (#1744) touches the same kernel but is not
a dependency. Combining the two opt-ins requires its own validation.

No site

Links install, modelos, releases.