Issues / #729

#729 KV int8 vs fp16 KLD compare at 240K context

closed · @jerry78424 · 2 comentários · No GitHub

BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentation

Descrição


The [2026-09-27 bench](https://github.com/Niko1221/Strata/blob/main/bench/results/2026-09-27-kv-q4/README.md) stopped at 8K and reported "int8 is indistinguishable from fp16". 

Two questions were left open:
1. Does that hold at long context?
2. "Indistinguishable" was measured teacher-forced. Does it survive **free generation**, where each arm's
   own tokens feed its own KV and a single flipped near-tie makes the two texts diverge?

## Setup
- IQ3_S on an RTX 5090 32 GB / 192 GB RAM, engine 0.1.34 and re-run on 0.1.38 (results are **bit-identical**
  between the two versions on this path - the update did not change numerics here).
- One 240,000-token document (15 repo docs + 6 large source files); 16 prefixes at 15K..240K.
- Arms differ only in `--kv fp16` vs `--kv int8`; fixed expert set (`--expert-cache 7000` budget,
  `--adapt-swaps 0`), `--prefill 8192`, `--kv-resident 0`, `--max-context 262144`.
- Measurement hygiene: each arm runs with fresh loaded model instance.

## Method (native packs constrain the probe)
Native IQ packs leave the per-token `--dump-logits` loop before the dump site (`--prefill 0` is refused;
69de404 made the empty file honest), so there is no per-position logit series. What exists instead:
- **Teacher-forced probe**: `STRATA_DUMP_FIRST_LOGITS` - one full-vocab (248,320 f32) row at the last prompt
  position, i.e. the next-token distribution after attending over the entire batched-prefill KV.
  Verified byte-identical across reruns (temp-0 greedy, `--spec 4` lossless).
- **Fork probe**: the same arm with `--max-new 32 --temperature 0` (generate mode defaults to temperature
  1.0 - pin it); the engine prints the generated token ids; the two arms' 32-token sequences are compared
  for the first differing token. Same prefill cost, so this is nearly free.
## Distribution: no growth with depth, heavy tail at near-ties
| Depth | KL(int8‖fp16) | JS | Same top-1 | Fork token (of 32) |
|---:|---:|---:|---|---:|
| 15K | 0.00026 | 5.5e-05 | yes | 22 |
| 30K | 0.0018 | 3.0e-04 | yes | 19 |
| 45K | **1.13** | 0.190 | **no** | 0 |
| 60K | 0.00002 | 6e-06 | yes | 8 |
| 75K | 0.0087 | 2.2e-03 | yes | 6 |
| 90K | 0.00008 | 2.1e-05 | yes | 15 |
| 105K | 0.00015 | 4.4e-05 | yes | 8 |
| 120K | 0.00092 | 2.0e-04 | yes | 15 |
| 135K | 0.047 | 8.5e-03 | yes | 4 |
| 150K | 0.0005 | 9.8e-05 | yes | identical |
| 165K | **0.36** | 0.074 | yes | 20 |
| 180K | 0.0019 | 4.4e-04 | yes | 6 |
| 195K | 0.00007 | 1.6e-05 | yes | 3 |
| 210K | 0.00005 | 1.4e-05 | yes | identical |
| 225K | 0.00006 | 1.8e-05 | yes | identical |
| 240K | **0.43** | 0.061 | yes | 10 |
- Median KL ~1e-3 and **no growth with depth**: several 195K-240K points sit at ~6e-5, below the 8K mean
  of 0.114. The sqrt(d) averaging of K/V quantization noise holds out to 240K.
- The three large points are high-entropy **positions**, not deep ones (45K flips the top-1 between two
  near-tied tokens; 165K has both arms at NLL > 10). Same heavy-tail shape as the 8K bench.

## Free generation: the texts fork quickly anyway
**13 of 16 depths fork within 32 greedy tokens** (indices 0,3,4,6,6,8,8,10,15,15,19,20,22; median ~8);
3 identical (150K, 210K, 225K). The sharpest point: **195K forks at token 3 with a teacher-forced KL of
7e-5** - the average distribution is nearly untouched, yet along the generated path some step always has
a top-2 gap smaller than the int8 noise (a few percent flip probability per token).

## What this does and does not say
- **Supported**: int8 KV's per-step distribution shift is tiny and does not accumulate with context;
  aggregate quality (the 8K perplexity numbers) is unaffected. int8 staying the default is right.
- **Not supported**: "the output is the same as fp16". At 240K scale a greedy continuation from the same
  prompt diverges within a handful of tokens at most positions. That is expected at near-ties (fp16 itself
  is a coin flip there).
- **Not measured**: whether the diverged continuations are quality-equivalent (task-level eval, e.g. a
  needle bench at 240K, is the missing piece; the 9-27 needles only showed retrieval survives q4_0).

## Limitations
One model (IQ3_S), one corpus, 16 points, one teacher-forced row per arm.
Fork indices bound the near-tie flip rate coarsely (n=32 steps per depth). Prompts were fed as raw token
ids (`--tokens-file`), so **no chat template (jinja) was applied** - the arms continue a raw token stream
with no template markers anywhere in the corpus, and the ChatML markers the model emitted in its outputs
are fully learned behavior, not echoes of the input. The degenerate outputs observed on both sides (an
endoftext repeat loop at 135K on the fp16 arm, a stray tool-call end at 75K) are partly artifacts of that
un-templated state and do not imply real-serving degeneracy. The probe needs
`STRATA_DUMP_FIRST_LOGITS` (a debug env hook); a first-class per-position dump path for native packs
would make this exact experiment much cheaper to repeat.

No site

Links install, modelos, releases.