Pull requests / #1744
#1744 CUDA SM86: opt-in IQ3 two-stage prefill build
open · @InB4DevOps · 0 comments · View on GitHub
Setup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Description
## Summary Add the off-by-default **`STRATA_CUDA_SM86_IQ3_STAGE2` CMake option**, requiring an explicit CUDA architecture 86 target. It builds a CUDA-only specialization; runtime opt-in remains `STRATA_PF_FUSED=1 STRATA_PF_IQ3_STAGE2=1`. Only **IQ3_S/IQ4_NL and IQ3_XXS/IQ4_NL** gate/up–down pairs use two activation stages. Every other pair and the default retain four stages. **Updated scope:** the shared kernel and C++ test edits from the initial draft have been fully reverted. The specialization is selected only in the CUDA CMake branch and only executes on runtime SM86. HIP's source/build branch and SYCL's build are unchanged, so the original missing-ROCm gate no longer applies to this patch. Other CUDA devices use the original implementation. SM86 includes other Ampere cards, but only the RTX 3060 was measured. Windows/other hardware remain untested. ## What changed - CUDA-only stage specialization with matching shared-memory sizing, copy-wait depth and launch attributes; conversion/routing helpers and the fallback are reused from the unchanged original implementation. The specialized kernel body must keep its arithmetic in sync with the original. - Instantiate the two-stage family only for the eligible kernel types; select it only when both formats of the layer match the measured pairs. - Fresh-process off/on numerical parity checks and scoped measurement evidence. The blanket two-stage experiment is **not** included: it regressed IQ2_S by 4.8–5.8% at some sizes. Those kernels retain four stages. Weight prefetch distance, tile/grid selection, residency, decode, setup defaults and HIP policy are unchanged. This branch is independently based on `fb58e0db`, with no prefetch/head experiments. ## Measured evidence RTX 3060 **12 GB / 100 W**, i7-12700KF AVX2, Linux, CUDA 13.4.92, native IQ3_XXS, explicit 4096-token prefill and static residency. Six balanced pairs per realistic prompt, 128 generated tokens: | Prompt | Input tokens | Median paired prompt-speed gain | Faster pairs | |---|---:|---:|---:| | Parser/test review | 12,247 | +0.199% | 5/6 | | Runtime documentation | 20,324 | +0.304% | 5/6 | | C++ prefill review | 30,757 | +0.239% | 5/6 | All 18 pairs had identical output IDs and draft counts and were repeatable. The clearer effect is in the IQ3 expert-layer workload: **+2.65–17.02%** across six format/size combinations, all 18 process-level shape pairs faster. Untouched IQ2 kernels stayed near baseline, rather than showing the global policy's regressions. These are **prototype measurements of the same selective policy**. Final isolated validation is separately recorded, not pooled into that series. The whole-engine effect is modest; A/A controls include roughly +/-1.5% individual outliers. This is a targeted kernel-tuning proposal, not a substantial general inference speedup. No occupancy increase was measured, and no combined-prefetch claim is made. ## Validation - CUDA SM86 Release: default-off reference build and this enabled build of `strata` and `prefill_fused_iq_test`. - Architecture-86 requirement and mutually exclusive build options checked. - Shared kernel/test bytes, HIP build branch and SYCL build file verified unchanged from upstream. - Six formats × both tiles: isolated off/on fingerprints and prior disabled-build comparison against the default-off build, retaining the existing FP64/MMQ checks. - Isolated two-pair checks on all three realistic prompts for token IDs, drafts and repeatability. - `git diff --check`. - Non-SM86 predicate checked at compile time; Windows/other GPU hardware runtime untested. Details: `docs/benchmarks/2026-10-09-rtx3060-iq3-stage2/README.md`. Current validation is in `cuda-only-evidence.json`/`cuda-only-parity.log`; earlier evidence remains separately recorded, with the exact public-source prompts. The separate one-superblock prefetch proposal (#1742) touches the same kernel but is not a dependency. Combining or merging both policies needs additional validation.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.