Pull requests / #293

#293 KV: optional Hadamard rotation for int8 K/V (STRATA_KV_ROT=1), with the drafter and batched verify made rotation-consistent

closed · @sergqwer · 0 comentários · No GitHub

NVIDIA / CUDAModels & quantsWindows

Descrição

## Summary

INT8 K and V can now go through the same Walsh-Hadamard rotation that `--kv q4_0` already uses (`kv_q4.hpp`). The
rotation spreads outlier channels over the 64-value scale groups before quantization. The queries are rotated too
(`<Hq, Hk> = <q, k>`) and the output is rotated back. It is **opt-in**: `STRATA_KV_ROT=1`.

Along the way, two places that were not rotation-consistent are fixed. They matter for Q4_0 today and for rotated INT8:

- `verify.cpp`: the batched window path (`qb`) never rotated its queries. That was harmless only because Q4_0 was
  kept off that path.
- `Prefill::draft_kv`: the MTP drafter's batched prompt pass rotated the prompt's K/V only for Q4_0, while the
  drafter's own decode (`mtp.cpp`) rotates by `kv_rot`. With rotated INT8 the drafter would score rotated queries
  against unrotated prompt keys. On the fork this cost 1.60 vs 2.40 tokens per round after an 8K prompt.

## Measured

Measured on current main (0.1.29, d6708a4) with ISTA-DASLab's IQ2_XS, RTX 5090 32 GB, Ryzen 9 9950X3D, 128 GB DDR5, Windows 11. KL is first-token KL divergence (`STRATA_DUMP_FIRST_LOGITS`, #276); two runs of main itself differ by 0 on 5 of 6 short prompts and by 0.00013 on the sixth. First-token KL to fp16 K/V:

| | 1K | 2K | 4K | 32K | mean |
| --- | ---: | ---: | ---: | ---: | ---: |
| IQ2_XS, int8 | 0.0018 | 0.0035 | 0.0001 | 0.0033 | 0.0022 |
| IQ2_XS, int8 rotated | 0.0036 | 0.0073 | 0.0003 | 0.0104 | 0.0054 |
| NVFP4 pack, int8 | 0.0033 | 0.0152 | 0.0001 | 0.0079 | 0.0066 |
| NVFP4 pack, int8 rotated | 0.0063 | 0.0079 | 0.00005 | 0.0063 | 0.0051 |

On the fork, on real K/V, the attention output's error against FP32 went from 0.306% to 0.265%. End to end it helps
the NVFP4 pack and hurts IQ2_XS, which is why it is off by default. I don't have an explanation for the IQ2_XS
result; the same code paths run for both packs. It may be worth a look before anyone turns it on by default.

## Correctness

Without the switch, the generated tokens are identical to main's (128 tokens, IQ2_XS). With `STRATA_KV_ROT=1`
drafting works (2.57 tokens per round).

## Switch

`STRATA_KV_ROT=1` rotates INT8 K/V.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

No site

Links install, modelos, releases.