Pull requests / #1742
#1742 CUDA SM86: opt-in one-superblock prefetch build
open · @InB4DevOps · 0 commentaires · Sur GitHub
BenchmarksSetup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Description
## Summary Add the off-by-default **`STRATA_CUDA_SM86_PREFETCH_ONE` CMake option**. With an explicit CUDA architecture 86 target, it builds a CUDA-only specialization for the existing native fused path. Runtime opt-in remains `STRATA_PF_FUSED=1 STRATA_PF_PREFETCH_ONE=1`. Gate/up weights are prefetched one superblock ahead instead of two. **Updated scope:** the initial draft edited shared CUDA/HIP source. That edit and the shared C++ test change have now been completely reverted. The new translation unit is selected only inside the CUDA CMake branch, with an additional runtime SM86 check. HIP's implementation/build branch and SYCL's build remain unchanged; there is no HIP change to build. Other CUDA devices use the original implementation. SM86 includes more cards than the 3060; measurements remain 3060-specific. Windows and other GPU hardware have not been tested. ## What changed - CUDA-only kernel specialization for gate/up, reusing conversion/routing helpers and the fallback by including the unchanged original implementation. Down keeps its original instance. The specialized body must retain the original arithmetic. - A fresh-process numerical parity wrapper for six format pairs and both tile sizes. - Scoped benchmark report, exact public-source prompts, per-pair timing/output-hash evidence, and final isolated validation results. No activation-stage, tile/grid, residency, decode, model/setup-default or HIP-kernel policy change. This branch is independently based on `fb58e0db` and contains none of the other experimental engine patches. ## Measured evidence RTX 3060 **12 GB / 100 W**, i7-12700KF AVX2, Linux, CUDA 13.4.92, native IQ3_XXS, explicit 4096-token prefill and static residency. Six balanced pairs per realistic prompt, 128 generated tokens: | Prompt | Input tokens | Median paired prompt-speed gain | Faster pairs | |---|---:|---:|---:| | Parser/test review | 12,247 | +0.548% | 5/6 | | Runtime documentation | 20,324 | +0.813% | 6/6 | | C++ prefill review | 30,757 | +0.802% | 6/6 | All 18 pairs had identical token IDs and draft counts, with within-arm repeatability. Three independent fixed-work process pairs per shape gave **+1.71–6.52%** across twelve format/size combinations; all 36 shape/process pairs were faster with matching hashes. These are **prototype measurements of the same policy**, not inflated to represent the separate MMQ-to-fused speedup. The isolated extraction's tests/smoke pairs are reported separately. A/A controls had near-zero medians but individual outliers up to about +/-1.5%; all outliers remain in the evidence. No broad GPU/model/default or quality-equivalence claim is made. ## Validation - CUDA SM86 Release: default-off build and opt-in build of `strata` and `prefill_fused_iq_test`. - Configuration rejects the opt-in without architecture 86 and rejects combining the two build variants. - Shared kernel/test files verified byte-identical to upstream; HIP CMake branch and SYCL build file unchanged. - Isolated off/on outputs: six formats × both tiles, plus prior disabled-build default-build fingerprint comparison; existing FP64/MMQ numerical checks retained. - Isolated two-pair whole-engine checks on the same three realistic prompts; token IDs, draft counts and repeatability checked. - `git diff --check`. - Non-SM86 dispatch predicate checked at compile time; other GPU hardware/Windows runtime untested. Details and commands: `docs/benchmarks/2026-10-09-rtx3060-prefetch-one/README.md`. Current validation: `cuda-only-evidence.json` and `cuda-only-parity.log`; historical prototype/initial extraction: `evidence.json` and `isolated-parity.log`. Exact prompts are included. The separate IQ3-only stage-buffering proposal (#1744) touches the same kernel but is not a dependency. Combining the two opt-ins requires its own validation.
Sur le site
Liens install, modèles, releases.