Pull requests / #1744

#1744 CUDA SM86: opt-in IQ3 two-stage prefill build

open · @InB4DevOps · 0 コメント · GitHub で見る

Setup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

本文

## Summary

Add the off-by-default **`STRATA_CUDA_SM86_IQ3_STAGE2` CMake option**, requiring an
explicit CUDA architecture 86 target. It builds a CUDA-only specialization; runtime
opt-in remains `STRATA_PF_FUSED=1 STRATA_PF_IQ3_STAGE2=1`. Only **IQ3_S/IQ4_NL and IQ3_XXS/IQ4_NL**
gate/up–down pairs use two activation stages. Every other pair and the default
retain four stages.

**Updated scope:** the shared kernel and C++ test edits from the initial draft have
been fully reverted. The specialization is selected only in the CUDA CMake branch
and only executes on runtime SM86. HIP's source/build branch and SYCL's build are
unchanged, so the original missing-ROCm gate no longer applies to this patch.
Other CUDA devices use the original implementation. SM86 includes other Ampere
cards, but only the RTX 3060 was measured. Windows/other hardware remain untested.

## What changed

- CUDA-only stage specialization with matching shared-memory sizing, copy-wait
  depth and launch attributes; conversion/routing helpers and the fallback are
  reused from the unchanged original implementation. The specialized kernel body
  must keep its arithmetic in sync with the original.
- Instantiate the two-stage family only for the eligible kernel types; select it
  only when both formats of the layer match the measured pairs.
- Fresh-process off/on numerical parity checks and scoped measurement evidence.

The blanket two-stage experiment is **not** included: it regressed IQ2_S by
4.8–5.8% at some sizes. Those kernels retain four stages. Weight prefetch distance,
tile/grid selection, residency, decode, setup defaults and HIP policy are unchanged.
This branch is independently based on `fb58e0db`, with no prefetch/head experiments.

## Measured evidence

RTX 3060 **12 GB / 100 W**, i7-12700KF AVX2, Linux, CUDA 13.4.92, native IQ3_XXS,
explicit 4096-token prefill and static residency. Six balanced pairs per realistic
prompt, 128 generated tokens:

| Prompt | Input tokens | Median paired prompt-speed gain | Faster pairs |
|---|---:|---:|---:|
| Parser/test review | 12,247 | +0.199% | 5/6 |
| Runtime documentation | 20,324 | +0.304% | 5/6 |
| C++ prefill review | 30,757 | +0.239% | 5/6 |

All 18 pairs had identical output IDs and draft counts and were repeatable.
The clearer effect is in the IQ3 expert-layer workload: **+2.65–17.02%** across
six format/size combinations, all 18 process-level shape pairs faster. Untouched
IQ2 kernels stayed near baseline, rather than showing the global policy's regressions.

These are **prototype measurements of the same selective policy**. Final isolated
validation is separately recorded, not pooled into that series. The whole-engine
effect is modest; A/A controls include roughly +/-1.5% individual outliers. This
is a targeted kernel-tuning proposal, not a substantial general inference speedup.
No occupancy increase was measured, and no combined-prefetch claim is made.

## Validation

- CUDA SM86 Release: default-off reference build and this enabled build of `strata` and `prefill_fused_iq_test`.
- Architecture-86 requirement and mutually exclusive build options checked.
- Shared kernel/test bytes, HIP build branch and SYCL build file verified unchanged from upstream.
- Six formats × both tiles: isolated off/on fingerprints and prior disabled-build
  comparison against the default-off build, retaining the existing FP64/MMQ checks.
- Isolated two-pair checks on all three realistic prompts for token IDs, drafts
  and repeatability.
- `git diff --check`.
- Non-SM86 predicate checked at compile time; Windows/other GPU hardware runtime untested.

Details: `docs/benchmarks/2026-10-09-rtx3060-iq3-stage2/README.md`. Current validation
is in `cuda-only-evidence.json`/`cuda-only-parity.log`; earlier evidence remains
separately recorded, with the exact public-source prompts.

The separate one-superblock prefetch proposal (#1742) touches the same kernel but is not
a dependency. Combining or merging both policies needs additional validation.

関連リンク

インストール・モデル・リリースへの站内リンク。