Pull requests / #258
#258 fused_gr: TILE=1280 kernel specialization for sm_75 (Turing) down-projection
closed · @hireymage · 0 commentaires · Sur GitHub
Description
Title: fused_gr: TILE=1280 kernel specialization for sm_75 (Turing) down-projection Body: ## What this does On Turing (sm_75) the fused gate/rescale down-projection kernel runs with the default tile geometry sized for larger shared-memory budgets. Turing has 64 KiB of shared memory per block (opt-in), which fits a different, more efficient shape: a TILE=1280 specialization fits **8 activation tokens in a single ~40 KiB launch** instead of the default tile's smaller working set, and drops register pressure (REG 76 → 64) enough to raise occupancy. - Adds a `sm_75`-guarded TILE=1280 instantiation of the down-projection kernel alongside the default geometry. - Bit-exact: the kernel still accumulates in strictly increasing (tile, chunk) order, so results are unchanged — verified greedy generation matches the default build with temperature 0. - No effect on other architectures (guarded, default path untouched). ## Measured on my machine i7-8700K (AVX2), RTX 2070 (sm_75), Qwen3.8-Flash-Next native Q2_0 pack (spec verified from the GGUF header, ~125.7B-parameter MoE), standalone kernel microbench, temperature 0. - Kernel level (verify-window timing of the down-projection launches, stage profiler on): - GDN down: **2.73–2.80 → 1.88–1.98 ms/window (−30 %)** - QSA down: **0.90–0.93 → 0.62–0.65 ms/window (−30 %)** - End to end on this rig: within noise, **because decode is CPU expert-pool bound here** (the CPU AVX2 expert pool dominates the window). This change frees GPU time the CPU cannot consume yet, so its end-to-end benefit should appear on rigs with a faster CPU or a larger model where the GPU down-projection is the bottleneck. - Profiling context (Nsight Compute, verify-window kernel capture): the grouped-expert kernels on sm_75 are L1TEX-pipe bound (~95 % of peak, DRAM ~22–33 %), driven by uncoalesced 8-byte weight reads. The tile change improves occupancy and per-launch batching without touching the access pattern or the summation order. ## Please test on other devices I can only measure one machine. This is a Turing-specific specialization, so other sm_75 cards (RTX 2060/2070/2080 variants) are the direct target: - **RTX 20-series owners**: please rerun a fixed-prompt greedy decode (temperature 0) before/after and compare wall-clock decode time, and ideally per-window kernel time if you use the stage profiler. - **Non-Turing GPUs** should be unaffected (guard), but confirmation is welcome — especially Ampere+ cards with the larger default smem budget, to confirm the default path is untouched. - Rigs with fast CPUs (where GPU decode share is larger) are where the end-to-end effect should actually show. Greedy outputs should be identical before/after; if you see a difference, that's a bug, please report it.
Sur le site
Liens install, modèles, releases.