Pull requests / #21
#21 kv: add Q4_0 KV cache mode with FWHT-256 Hadamard rotation
closed · merged 2026-09-27 · @code-martin · 0 Kommentare · Auf GitHub
Server & APINVIDIA / CUDAModels & quants
Beschreibung
## What
Adds an optional 4-bit (`Q4_0`) quantized KV cache mode with Fast Walsh-Hadamard Transform (FWHT-256) rotation to Strata:
- Activated via `--kv q4_0` flag in the engine or server configuration.
- Halves KV cache memory footprint compared to `Q8_0` (down to ~4.5 bytes per token per layer per head), unlocking double context capacity on memory-constrained systems (e.g. 256k context).
- Implements fast in-register FWHT-256 rotation kernel before quantization to eliminate activation outliers across the head dimension ($D=256$).
- Maintains attention score invariance $\langle Hq, Hk \rangle = \langle q, k \rangle$ by applying the same orthogonal Hadamard matrix to query tokens.
- Fully compatible with speculative verification (`verify.cpp`), multi-token prediction (`mtp.cpp`), and QSA decode streaming (`qsa_decode_attn.cu`).
## Why
At 256k context, unquantized or 8-bit KV caches consume substantial memory. Standard 4-bit quantization on raw key-value vectors often suffers quality loss due to channel outliers. Applying an orthogonal Walsh-Hadamard transform distributes outlier energy uniformly across all 256 dimensions before block quantization, preserving perplexity with zero score drift.
## How
1. **FWHT-256 Kernel (`src/kernels/cuda/kv_q4.cu`, `include/strata/kernels/kv_q4.hpp`)**:
- In-register butterfly network across 8 warps per block with bitwise XOR shuffles.
- Normalization factor $1/\sqrt{256} = 1/16$.
2. **Quantized Key/Value Store & Decode Attention (`src/kernels/cuda/qsa_decode_attn.cu`)**:
- `Q4_0` block layout: 1 FP16 scale + 16 bytes for 32 int4 nibbles.
- Vectorized dequantization in warp registers during flash-attention score accumulation.
3. **Parity & Correctness Test (`src/kernels/kv_q4_parity.cpp`)**:
- Verifies Hadamard self-inverse $H \cdot H \cdot x == x$.
- Verifies dot-product conservation $\langle Hq, Hk \rangle == \langle q, k \rangle$.
- Verifies quant/dequant parity and max error bounds.
4. **Tooling & Build**:
- Added `kv_q4_parity` target in `CMakeLists.txt`.
- Fixed path resolution in `tools/_paths.py`, `tools/pack_layer.py`, `tools/strata_pack.py` when repo is placed near filesystem roots.
## Verification
- `kv_q4_parity.exe` passes 100% on CUDA (FWHT-256 bitwise match, attention score invariance, block bounds).
- Tested end-to-end on Qwen3.8-Flash-Next 125B (swift-1.5-iq2_xs) with 256k context window via OpenAI /v1/chat/completions API.Mehr auf der Site
Links zu Install, Modellen, Releases.