Pull requests / #343
#343 Don't pursue: BF16 decode GEMV via Turing tensor cores (sm_75) - measured negative
closed · @hireymage · 0 comentários · No GitHub
BenchmarksNVIDIA / CUDAModels & quantsDocumentation
Descrição
**Don't pursue: BF16 decode GEMV via Turing tensor cores (sm_75) — measured negative** Question: the QSA prompt attention rides the Turing tensor cores (#270-style `mma.sync.m16n8k8`); can the decode GEMVs (`bf16_gemv.cu`) do the same? This PR adds an isolated benchmark comparing the production warp kernel against the tensor-core path on the real decode shapes (n_in = 2560, BF16 weights), with three TC variants covering the conversion cost from every angle (W preconverted once / in-register / fused no-reduce form) and a full k-split occupancy sweep for the TC side. No inference-path changes. **Result: the TC variant loses on every real shape, 1.3–2.2×, even at its best config** — after conversion overhead: | Shape | Baseline µs | Best TC µs | Speedup | | --- | ---: | ---: | ---: | | ssm_alpha/beta [2560,48] | 8.73 | 11.32 | 0.77× | | indexer.k_proj [2560,128] | 7.14 | 12.09 | 0.59× | | q_proj/router [2560,512] | 9.48 | 15.78 | 0.60× | | ple_value [2560,2560] | 35.08 | 77.73 | 0.45× | Parity passes (max err vs fp64 ref at fp16-accumulation level, 1.7e-6–7.2e-6; BF16→FP16 conversion exact, 0/2,000,000). The warp kernel is already bandwidth-bound at 374 GB/s; the TC path reaches only ~169 GB/s and a single-token GEMV throws away 15/16 of the m16 MMA tile. Measured on i7-8700K + RTX 2070 8 GB (sm_75, 36 SMs), engine 0.1.27 tree (`bf16_gemv.cu` unchanged through 0.1.30), qwen3.8-flash-next ~125.7B MoE Q2_0 in production. **Conclusions: keep the warp kernel; the benchmark is checked in so nobody has to measure this twice** and can rerun it on future hardware (native bf16, denser GEMV batching). Details and rerun instructions in `bench/results/2026-09-30-tc-gemv-sm75/README.md`. Limit: latency from back-to-back CUDA-event timing (no ncu profile); production engine ran concurrently on the same GPU.
No site
Links install, modelos, releases.