Pull requests / #21

#21 kv: add Q4_0 KV cache mode with FWHT-256 Hadamard rotation

closed · merged 2026-09-27 · @code-martin · 0 コメント · GitHub で見る

Server & APINVIDIA / CUDAModels & quants

本文

## What

Adds an optional 4-bit (`Q4_0`) quantized KV cache mode with Fast Walsh-Hadamard Transform (FWHT-256) rotation to Strata:
- Activated via `--kv q4_0` flag in the engine or server configuration.
- Halves KV cache memory footprint compared to `Q8_0` (down to ~4.5 bytes per token per layer per head), unlocking double context capacity on memory-constrained systems (e.g. 256k context).
- Implements fast in-register FWHT-256 rotation kernel before quantization to eliminate activation outliers across the head dimension ($D=256$).
- Maintains attention score invariance $\langle Hq, Hk \rangle = \langle q, k \rangle$ by applying the same orthogonal Hadamard matrix to query tokens.
- Fully compatible with speculative verification (`verify.cpp`), multi-token prediction (`mtp.cpp`), and QSA decode streaming (`qsa_decode_attn.cu`).

## Why

At 256k context, unquantized or 8-bit KV caches consume substantial memory. Standard 4-bit quantization on raw key-value vectors often suffers quality loss due to channel outliers. Applying an orthogonal Walsh-Hadamard transform distributes outlier energy uniformly across all 256 dimensions before block quantization, preserving perplexity with zero score drift.

## How

1. **FWHT-256 Kernel (`src/kernels/cuda/kv_q4.cu`, `include/strata/kernels/kv_q4.hpp`)**:
   - In-register butterfly network across 8 warps per block with bitwise XOR shuffles.
   - Normalization factor $1/\sqrt{256} = 1/16$.
2. **Quantized Key/Value Store & Decode Attention (`src/kernels/cuda/qsa_decode_attn.cu`)**:
   - `Q4_0` block layout: 1 FP16 scale + 16 bytes for 32 int4 nibbles.
   - Vectorized dequantization in warp registers during flash-attention score accumulation.
3. **Parity & Correctness Test (`src/kernels/kv_q4_parity.cpp`)**:
   - Verifies Hadamard self-inverse $H \cdot H \cdot x == x$.
   - Verifies dot-product conservation $\langle Hq, Hk \rangle == \langle q, k \rangle$.
   - Verifies quant/dequant parity and max error bounds.
4. **Tooling & Build**:
   - Added `kv_q4_parity` target in `CMakeLists.txt`.
   - Fixed path resolution in `tools/_paths.py`, `tools/pack_layer.py`, `tools/strata_pack.py` when repo is placed near filesystem roots.

## Verification

- `kv_q4_parity.exe` passes 100% on CUDA (FWHT-256 bitwise match, attention score invariance, block bounds).
- Tested end-to-end on Qwen3.8-Flash-Next 125B (swift-1.5-iq2_xs) with 256k context window via OpenAI /v1/chat/completions API.

関連リンク

インストール・モデル・リリースへの站内リンク。