Pull requests / #540

#540 volta (sm_70) and turing (sm_75): faster prompt reading - BF16 projections on the FP16 tensor cores, and a faster pre-Turing attention kernel

closed · @sskver · 0 comentarios · En GitHub

BenchmarksNVIDIA / CUDAWindows

Descripción

i got the experimental sm_70 build running on my Tesla V100 32GB (IQ2_XS) and found two things that were slow there, so here are fixes for both. i know volta isn't supported, this just makes the existing experimental build faster

it's the V100 build you mentioned in #101 / #130

### what changes

1. **BF16 GEMMs on volta and turing.** these cards have no BF16 tensor cores, so cuBLAS runs them as a CUDA-core kernel (`magma_sgemmEx`), about 10% of the GPU time in a prompt on a V100. on cc 7.0 and 7.5 i now convert the operands to FP16 in the existing scratch buffer and use the FP16 tensor-core path. `STRATA_BF16_VIA_F16=0` turns it off

2. **the attention kernel.** on sm_70 the tensor-core prompt attention declines, so `attn_chunk_kernel` does the work (about 20% of a 20K prefill). i cut its shuffles and shared memory traffic. this one is bit-exact, same operands in the same order. (#600 adds a tensor-core prompt attention for sm_70, which would take over prefill there. this kernel still runs decode and the verify windows, and the two PRs share no files)

### numbers

V100, prompt processing speed, v0.1.37, one run each, `--kv int8`:

| prompt | upstream | + attention fix | + attention and BF16 fix |
|---|---|---|---|
| 8K tokens | 874 tok/s | 934 (+7%) | 1066 (+22%) |
| 20K tokens | 896 tok/s | 960 (+7%) | 1100 (+23%) |

decode doesn't change (64.5 vs 66.0 tok/s, within noise)

**turing:** @fks ran the BF16 change on a Quadro RTX 4000 (8 GB, sm_75, CUDA 12.8, same binary with `STRATA_BF16_VIA_F16=0` vs default):

| prompt | off | on |
|---|---|---|
| 3.7K tokens | 131.1 tok/s | 141.8 (+8%) |
| 7.2K tokens | 137.0 tok/s | 149.5 (+9%) |

in nsys the CUDA-core BF16 kernel was 31.8% of GPU kernel time and is gone with the path on, total kernel time 9.18 s -> 6.73 s. the wall-clock gain is smaller because that card streams most experts over PCIe. one run per cell, and exactness of the conversion wasn't measured on sm_75

### accuracy

the attention commit alone gives byte-identical first-window logits to upstream at 8K and 20K.

the BF16 one is not bit-identical, since FP16 accumulation rounds differently and tiny values lose precision. i compared 24 different slices of source code, 1.5K to 5.5K tokens each, and the next-token distribution after prefill: mean KL 8.4e-3, max 3.7e-2, same top-1 in 23 of 24 (the odd one is a near tie), and perplexity of the true next token 3.448 vs 3.420, so no measurable change. that's source code only, on the V100, and only right after the prompt, not a long generation. the BF16X2 remainder GEMMs stay on cuBLAS because their values are too small for FP16

### tests

the `*_parity` ctests have the same results as plain upstream on my V100 (29 pass, same 6 fail, all failing before my change too: needs missing data, AVX2 on my CPU, tensor-core attention declining on sm_70)

### caveats

i could only test sm_70 myself. the sm_75 numbers are from @fks on one 8 GB card, one run per cell, without an exactness check of the conversion, and my accuracy numbers above are V100 only. the attention kernel also runs decode on every GPU. the arithmetic is the same, but it uses more registers (80 vs 38 on sm_70), so if you'd rather i limit it to pre-sm_75 i'm happy to. the speedup numbers are single runs and move a few percent between sessions

En el sitio

Enlaces a install, modelos, releases.