Pull requests / #741

#741 prefill: BF16 projections on FP16 tensor cores below sm_80 (+8-17% prefill on a 2080 Ti)

closed · @konijiwa110 · 0 Kommentare · Auf GitHub

BenchmarksAMD / HIPNVIDIA / CUDAModels & quants

Beschreibung

## What

Cards below sm_80 (RTX 20, GTX 16, V100) have FP16 tensor cores but none for BF16, so cuBLAS runs the prefill's BF16
GEMMs (`Gemm::bf16`: hyper-connection down/up, inject, SSM alpha/beta, indexer projections) on the CUDA cores. On an
RTX 2080 Ti at those shapes (T = 4096) that is 6-7 TFLOPS in BF16 against 34-48 TFLOPS in FP16.

There, `Gemm::bf16` now converts W and X to FP16 in the existing dequantization scratch (W at its start, X in slices of
at least 256 rows) and calls `cublasGemmEx` in FP16 with the same FP32 accumulation:

- a BF16 value converts exactly when its magnitude is in [2^-14, 65504]; smaller ones round to an FP16 subnormal
  (absolute error <= 2^-25), larger ones saturate to +-65504 and are counted (`STRATA_PREFILL_BF16_F16_CHECK=1` prints
  the count; none in the prompts measured below)
- a scratch too small for W + 256 rows keeps the BF16 GEMM
- HIP builds are unchanged

`STRATA_PREFILL_BF16_F16=0` keeps the BF16 GEMM; `=1` takes the FP16 path on any card (used by the test).

New test `prefill_bf16_gemm_test` (CUDA, skips without a device): the call shapes above incl. a sliced X, `beta = 1`
accumulation into a strided `ldy` with an offset, and the scratch-too-small fallback, against an FP64 reference
(tolerance 1e-5), and cells outside the product left untouched.

## Measured

Swift 1.5 IQ3_XXS on an RTX 2080 Ti 22 GB, same build A/B with `STRATA_PREFILL_BF16_F16=0` / unset, prefill tok/s
after warm-up:

| prompt | BF16 | FP16 path |
|---|---|---|
| 4K | 815 | 877 (+8%) |
| 16K | 933 | 1089 (+17%) |
| 32K | 976 | 1126 (+15%) |
| 128K | 848 | 973 (+15%) |

The `hc read` phase (`STRATA_PREFILL_TIMING`) at 128K: 29.8 s -> 13.8 s. Decode unchanged. No saturation in any of
these prompts. Greedy output differs from the BF16 path after a few dozen characters, which is the same level as two
runs of the unchanged engine differ from each other (16-105 characters); answer quality looked the same.

`prefill_bf16_gemm_test`: all cases PASS, max error 4e-8 to 2.8e-7.

Only tested on sm_75 (2080 Ti); cards at sm_80+ never take the path unless the env forces it.

Developed with an AI coding assistant; all numbers measured on an RTX 2080 Ti 22 GB / i5-12490F / 64 GB.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Mehr auf der Site

Links zu Install, Modellen, Releases.