Pull requests / #452
#452 qsa_prompt_attn: Q4_0 KV on tensor cores (mode 4) - long prompts with --kv q4_0 ~19% faster
closed · @architectds · 0 comentarios · En GitHub
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux
Descripción
## What The tensor-core prompt attention (D-1, `qsa_prompt_attn_batch`) handles FP16, INT8 and K8V4 KV but refuses Q4_0. With `--kv q4_0`, every prompt's QSA attention therefore ran on the FP32 per-query kernel (`qsa_decode_attn_batch`, 32 queries at a time). This adds a mode for Q4_0 K and V. - **Mode 4 is mode 1 (int8) with a scale per 32 values.** Each q4_0 block's 16 code bytes become 32 exact int8 codes (code - 8) at gather, and its fp16 scale goes into the per-group scales: 8 groups of 32 per row, where int8 has 4 of 64. The scales are applied in FP32 as int8's are: q.k per 32-dim group, and for p.v each group's scale is folded into p (two groups per warp). - **Signed scales.** ggml's q4_0 scale is `d = max / -8`, so about half of a real pool's blocks have a negative scale. The p.v fold is bounded by the chunk's largest `|scale|`. - **Everything else is mode 1's:** q and p split into hi + lo halves, the online softmax. The caller's rotation of q and of the output is unchanged. - **A/B switch.** `STRATA_PROMPT_ATTN_Q4=0` restores the old kernel. - **Test.** `qsa_prompt_attn_parity` gets Q4_0 cases (32K and 1.5K contexts) against FP64 and the old kernel. Their scales are signed and spread over two decades, as a pool's are. ## What it speeds up, and what it does not - **Prompts with `--kv q4_0`: +18.5% at 8K tokens, +18.8% at 16K, +19.9% at 30K.** With the old kernel the attention was 21-22% of a long prompt's GPU time on this machine. - **Other KV formats: unchanged.** Their dispatch is untouched, and the parity test's int8 and fp16 cases pass with the same numbers as before. - **Decode: unchanged.** This is the prompt path's batch attention only. - **HIP: unchanged.** The tensor-core kernels are CUDA-only, and HIP keeps the old kernel. ## Measured **Setup:** - RTX 5070 Ti 16 GB on PCIe 3.0 x16 (X370), Ryzen 9 5900XT, DDR4-2133, Windows 11, CUDA 13.0, built on v0.1.34. - Qwen3.8-Flash-Next IQ3_XXS, `--kv q4_0 --kv-resident 32768`, text-only args (4,560 expert slots), `--prefill auto` (8,192-token chunks). **Kernel** (`qsa_prompt_attn_parity 32768 2048 5`): | Case | Old kernel | Mode 4 | Speed | vs FP64 (old / new) | New vs old | | --- | ---: | ---: | ---: | --- | --- | | Q4_0, 32K context, 2,048 queries | 32.4 ms | 8.6 ms | 3.76x | 4.8e-6 / 6.2e-6 | 1.6e-6 of the output scale | | Q4_0, 1.5K context, 1,500 queries | 9.96 ms | 2.62 ms | 3.79x | 6.4e-6 / 5.1e-6 | 1.3e-6 of the output scale | The int8 and fp16 cases still pass. **Prompts:** - One binary, the old kernel (`STRATA_PROMPT_ATTN_Q4=0`) and mode 4 in alternating engine runs (old, new, old, new). - The display was off: on this card the desktop compositor otherwise takes up to a third of the GPU. - Every prompt has its own first line and corpus offset. - The binary also held three independent prompt-path changes (#439's batched gather and two others), switched off in both arms by their own variables, so the arms are v0.1.34 and v0.1.34 + this commit. | Prompt tokens | Old kernel (tok/s, run 1 / run 2) | Mode 4 (tok/s, run 1 / run 2) | Change (mean) | | ---: | ---: | ---: | ---: | | ~8K | 1,579 / 1,610 | 1,970 / 1,805 | +18.5% | | ~16K | 1,599 / 1,556 | 1,729 / 2,047 | +18.8% | | ~30K | 1,574 / 1,596 | 1,962 / 1,839 | +19.9% | **Quality** (not bitwise, as D-1 is not): - **Needles:** `tools/needle_bench.py`'s haystack and needle, driven through the engine. 8K, 16K and 32K at depths 10/30/50/70/90: 15 of 15 with the old kernel, 15 of 15 with mode 4. - **Greedy continuations** part from the old kernel's after a few tokens, as the int8 tensor-core kernel's do: the model amplifies any FP32-level change. - **A bug the tests caught.** An earlier draft bounded the p.v fold by the largest *positive* scale, as int8's are bounded. On real K/V that overflowed FP16 and the model returned NaN (needles 0 of 15). The parity test had passed it, because its synthetic scales were all positive. The test's scales are now signed, and it fails on that bound (new vs old 0.93 of the output scale). ## Not tested here - **HIP:** the kernels are CUDA-only, and HIP keeps the old kernel as before. - **Turing (sm_75):** mode 4 is in the v1 template kernel, which runs there through pairs of m16n8k8 MMAs. Not run. - A layer split, Linux, and packs other than IQ3_XXS (the attention does not depend on the pack). Developed with an AI coding assistant; all numbers measured on the machine above. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
En el sitio
Enlaces a install, modelos, releases.