Pull requests / #593
#593 prefill: run BF16 GEMMs through FP16 tensor cores on sm_70/sm_75
closed · @fks · 0 评论 · 在 GitHub 查看
BenchmarksNVIDIA / CUDAModels & quants
描述
Cards without BF16 tensor cores (Volta, Turing) ran every BF16 projection (hyper-connection, router, indexer, ...) as an FP32 SIMT kernel in cuBLAS (`magma_sgemmEx`): 17% of the GPU time of a 28,650-token prompt on a V100 (nsys, 0.1.35). Related: #139, which addressed the same kernels with a separate SGEMM. `Gemm::bf16` now converts W (into the existing dequantization scratch) and X (into an FP16 copy) to FP16 and takes the tensor-core `f16()` path. The conversion is exact for every BF16 value inside FP16's range; on that prompt no value overflowed, and the values below 2^-14 have an absolute error of at most 2^-25. The X copy (T x the widest K: 42 MB at T = 2048, 168 MB at the 8,192-token chunk the server used here) is carved from the prompt path's own buffers in `carve()` / `bytes_needed()`, like its other buffers, and only on cards that need it, so nothing is allocated after init. If it does not fit, or on a card with BF16 tensor cores, `bf16()` keeps the `cublasGemmEx` BF16 path. One V100, same binary, greedy output identical: 28,650-token prompt read in 26.9 s -> 23.8 s (1,063 -> 1,205 tok/s); 7,194 tokens +15%, 114,338 tokens +13%. `STRATA_BF16_VIA_F16=0|1` forces the old/new path; `STRATA_BF16_VIA_F16_CHECK=1` counts the values FP16 could not hold and prints them at exit. Not covered: sm_75 takes the same path but was not measured; only the Unsloth IQ4_XS pack was run. Written and measured with Claude Sonnet 5.5.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。