Pull requests / #1022
#1022 prefill/gdn: the Volta prompt recurrence chain-split (+7-15%) and the opt-in chunked path
closed · @ATIVX928 · 0 commentaires · Sur GitHub
Description
## What The prompt path's GDN recurrence for Volta, split out of #627: - the chain-split recurrence (four accumulators per thread, sm_70 only; FP32-level, the bit-exact variant is kept) - the chunked FlashQLA path as an opt-in study (`STRATA_GDN_CHUNK=1`/`=2`); it is 4-7x slower than the recurrence and stays off - `gdn_chunk_parity` and `gdn_chunk_bench` The upstream key-head / pipelined kernels are untouched; the Volta code was already re-based onto 0.1.38's key-head rewrite (the resolution from the aggregate v100-opt branch). ## Gating The Volta kernels are compiled under `STRATA_EXPERIMENTAL_SM60 && __CUDA_ARCH__ == 700`, and the dispatch asks for runtime cc 7.0, so the ready-made engine and other architectures are unchanged. The chunked path additionally needs `STRATA_GDN_CHUNK` to be set. ## Measured V100-SXM2-16GB, CUDA 12.8, `gdn_chunk_bench`, vs main's pipelined recurrence: | T | pipelined (main) | chain-split | gain | | ---: | ---: | ---: | ---: | | 1,024 | 1.027 ms | 0.893 ms | +15% | | 4,096 | 3.706 ms | 3.471 ms | +6.8% | | 8,192 | 7.481 ms | 6.920 ms | +8.1% | The chunked paths (FMA / wmma) are 4-7x slower at these shapes and remain opt-in. ## Tests - `gdn_chunk_parity`: 0 failures at T=64/128/200/320. Warp recurrence vs pipelined is bit-exact; the chain-split matches the recurrence (y max abs 3.7e-9, state 3.0e-8); the chunked wmma path is within its documented bounds (y 9.2e-6 of 6e-3). ## 32K end-to-end Same setup: single-card prefill 1546.0 and dual-card 2205.1 vs main's 1557.1 / 2207.0 - within the +-1% session noise. The recurrence is about 30 ms of a 21-second 32K prefill, so a 7-15% kernel gain cannot show end to end.
Sur le site
Liens install, modèles, releases.