Pull requests / #741
#741 prefill: BF16 projections on FP16 tensor cores below sm_80 (+8-17% prefill on a 2080 Ti)
closed · @konijiwa110 · 0 comentarios · En GitHub
BenchmarksAMD / HIPNVIDIA / CUDAModels & quants
Descripción
## What Cards below sm_80 (RTX 20, GTX 16, V100) have FP16 tensor cores but none for BF16, so cuBLAS runs the prefill's BF16 GEMMs (`Gemm::bf16`: hyper-connection down/up, inject, SSM alpha/beta, indexer projections) on the CUDA cores. On an RTX 2080 Ti at those shapes (T = 4096) that is 6-7 TFLOPS in BF16 against 34-48 TFLOPS in FP16. There, `Gemm::bf16` now converts W and X to FP16 in the existing dequantization scratch (W at its start, X in slices of at least 256 rows) and calls `cublasGemmEx` in FP16 with the same FP32 accumulation: - a BF16 value converts exactly when its magnitude is in [2^-14, 65504]; smaller ones round to an FP16 subnormal (absolute error <= 2^-25), larger ones saturate to +-65504 and are counted (`STRATA_PREFILL_BF16_F16_CHECK=1` prints the count; none in the prompts measured below) - a scratch too small for W + 256 rows keeps the BF16 GEMM - HIP builds are unchanged `STRATA_PREFILL_BF16_F16=0` keeps the BF16 GEMM; `=1` takes the FP16 path on any card (used by the test). New test `prefill_bf16_gemm_test` (CUDA, skips without a device): the call shapes above incl. a sliced X, `beta = 1` accumulation into a strided `ldy` with an offset, and the scratch-too-small fallback, against an FP64 reference (tolerance 1e-5), and cells outside the product left untouched. ## Measured Swift 1.5 IQ3_XXS on an RTX 2080 Ti 22 GB, same build A/B with `STRATA_PREFILL_BF16_F16=0` / unset, prefill tok/s after warm-up: | prompt | BF16 | FP16 path | |---|---|---| | 4K | 815 | 877 (+8%) | | 16K | 933 | 1089 (+17%) | | 32K | 976 | 1126 (+15%) | | 128K | 848 | 973 (+15%) | The `hc read` phase (`STRATA_PREFILL_TIMING`) at 128K: 29.8 s -> 13.8 s. Decode unchanged. No saturation in any of these prompts. Greedy output differs from the BF16 path after a few dozen characters, which is the same level as two runs of the unchanged engine differ from each other (16-105 characters); answer quality looked the same. `prefill_bf16_gemm_test`: all cases PASS, max error 4e-8 to 2.8e-7. Only tested on sm_75 (2080 Ti); cards at sm_80+ never take the path unless the env forces it. Developed with an AI coding assistant; all numbers measured on an RTX 2080 Ti 22 GB / i5-12490F / 64 GB. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
En el sitio
Enlaces a install, modelos, releases.