Pull requests / #258

#258 fused_gr: TILE=1280 kernel specialization for sm_75 (Turing) down-projection

closed · @hireymage · 0 comentarios · En GitHub

NVIDIA / CUDAModels & quants

Descripción

Title:
fused_gr: TILE=1280 kernel specialization for sm_75 (Turing) down-projection

Body:

## What this does

On Turing (sm_75) the fused gate/rescale down-projection kernel runs with
the default tile geometry sized for larger shared-memory budgets. Turing
has 64 KiB of shared memory per block (opt-in), which fits a different,
more efficient shape: a TILE=1280 specialization fits **8 activation
tokens in a single ~40 KiB launch** instead of the default tile's
smaller working set, and drops register pressure (REG 76 → 64) enough to
raise occupancy.

- Adds a `sm_75`-guarded TILE=1280 instantiation of the down-projection
  kernel alongside the default geometry.
- Bit-exact: the kernel still accumulates in strictly increasing
  (tile, chunk) order, so results are unchanged — verified greedy
  generation matches the default build with temperature 0.
- No effect on other architectures (guarded, default path untouched).

## Measured on my machine

i7-8700K (AVX2), RTX 2070 (sm_75), Qwen3.8-Flash-Next native Q2_0 pack
(spec verified from the GGUF header, ~125.7B-parameter MoE), standalone
kernel microbench, temperature 0.

- Kernel level (verify-window timing of the down-projection launches,
  stage profiler on):
  - GDN down: **2.73–2.80 → 1.88–1.98 ms/window (−30 %)**
  - QSA down: **0.90–0.93 → 0.62–0.65 ms/window (−30 %)**
- End to end on this rig: within noise, **because decode is CPU
  expert-pool bound here** (the CPU AVX2 expert pool dominates the
  window). This change frees GPU time the CPU cannot consume yet, so its
  end-to-end benefit should appear on rigs with a faster CPU or a larger
  model where the GPU down-projection is the bottleneck.
- Profiling context (Nsight Compute, verify-window kernel capture): the
  grouped-expert kernels on sm_75 are L1TEX-pipe bound (~95 % of peak,
  DRAM ~22–33 %), driven by uncoalesced 8-byte weight reads. The tile
  change improves occupancy and per-launch batching without touching the
  access pattern or the summation order.

## Please test on other devices

I can only measure one machine. This is a Turing-specific specialization,
so other sm_75 cards (RTX 2060/2070/2080 variants) are the direct target:

- **RTX 20-series owners**: please rerun a fixed-prompt greedy decode
  (temperature 0) before/after and compare wall-clock decode time, and
  ideally per-window kernel time if you use the stage profiler.
- **Non-Turing GPUs** should be unaffected (guard), but confirmation is
  welcome — especially Ampere+ cards with the larger default smem budget,
  to confirm the default path is untouched.
- Rigs with fast CPUs (where GPU decode share is larger) are where the
  end-to-end effect should actually show.

Greedy outputs should be identical before/after; if you see a difference,
that's a bug, please report it.

En el sitio

Enlaces a install, modelos, releases.