Issues / #1135

#1135 k8v4: deterministic token repetition on sequential generation (RTX 5090, 0.1.40) — byte-identical across two quants

closed · @gravitomagnetic · 4 commentaires · Sur GitHub

Setup & installNVIDIA / CUDAModels & quantsDocumentation

Description

Follow-up to my field report in #711 (comment [6016065570](https://github.com/Niko1221/Strata/pull/711#issuecomment-6016065570)) — filing as an issue so it can be tracked.

**Setup:** RTX 5090 32 GB, Strata 0.1.40 source build, Qwen3.8-Flash-Next MoE (12 QSA layers), 1M ctx YaRN×4, `--kv-resident 32768`, `--mtp --spec 4`, `--batch 2 --batch-mtp`, vision on, driver 595.91.07, CUDA 13.4. A/B: `--kv int8` vs `--kv k8v4`, fresh engine start per leg, temperature 0, `enable_thinking: false`.

**Symptom:** under k8v4, the decode probe "Count from 1 to 300, one number per line" repeats each integer:

```
int8:  1 2 3 4 5 6 7 8 9 10 11 12 13 …
k8v4:  1 2 2 3 3 4 4 5 5 6 6 7 7 8 8 9 9 10 10 10 10 11 11 11 11 12 12 12 12 13 13 13 1…
```

Runs of four appear around 10–13. Deterministic at temp 0.

**The sharpest clue:** the corrupted output is **byte-identical between IQ3_S and Q2_0** — two substrates sharing nothing about their weights (1.94 vs 1.32 MiB experts, different quantization). The only shared component is the k8v4 KV path, so the repetition is introduced by the value quantization/dequant itself, not by model weights or random distortion. That signature reads like a systematic grid/scale bias rather than inherent 4-bit noise — potentially fixable.

**What is NOT affected:** needle-in-haystack exact recall at ~200K tokens passes on both legs (keys stay INT8-exact, retrieval survives). Cold prefill is ~flat (−3%). Decode is actually **+30–33%** (98→130 IQ3_S, 151→197 Q2_0). So this is specifically short-range sequential generation — plausibly the recency signal in the value payload of the most recent cells ("what did I just emit").

**Repro:** the harness is in my last comment on #711; probes are cold-prefill / counting / needle. Single run per leg; the artifact reproduced identically in both quants' runs, so it is not a one-off.

**Question for you:** is this the expected distortion of rotated Q4_0 values at 32,768 resident cells, or worth a look at the dequant scale rounding? If expected, a line in the k8v4 docs ("avoid for enumeration/structured-output seats; retrieval unaffected") would help users choose.

Not tested yet: whether code, JSON, or other structured outputs show the same doubling.

Sur le site

Liens install, modèles, releases.