Issues / #1188

#1188 --kv k8v4 produces degenerate repetition when --kv-resident streams (0.1.40)

closed · @baiqvesse · 3 コメント · GitHub で見る

BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows

本文

## Summary

On 0.1.40, `--kv k8v4` combined with `--kv-resident` (the combination #711 / #705 newly enabled, and which the 0.1.40 notes say setup now offers) makes the model degenerate into repetition during free generation. The k8v4 format on its own is fine, and the failure is silent: no error or warning in the engine log, and long-context retrieval still works.

## Environment

| | |
|---|---|
| Strata | 0.1.40 (source + official `strata-windows-x64.zip`, `BUILD.json`: `archs [75,86,89,120]`, `ptx true`, `cuda 13.0`) |
| OS | Windows 11 Pro Insider Preview, build 29648 (10.0.29648) |
| GPU | NVIDIA GeForce RTX 5080, driver 617.14, sm_120, 16303 MiB |
| CPU / RAM | AMD Ryzen 9 9950X (16 cores) / 61.6 GB |
| Models | IQ3_XXS and stock IQ2_XS (both 2 shards) |

## How to reproduce

Config args: `--kv k8v4 --kv-resident 16384 --max-context 262144` (the engine clamps the requested 16384 up to its documented 20480 minimum — that is expected, and the log confirms it).

Request, greedy so it is deterministic:

```
POST http://127.0.0.1:8096/v1/chat/completions
{"model":"strata",
 "messages":[{"role":"user","content":"Write the numbers from 1 to 60 in words, separated by commas, with no other text."}],
 "max_tokens":400,"temperature":0,"reasoning_effort":"none"}
```

## Actual vs expected

**Expected** — what `int8`, `q4_0`, and `k8v4` *without* streaming all produce, byte for byte identical, ending naturally at 156 tokens:

```
one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen,
fourteen, fifteen, sixteen, seventeen, eighteen, nineteen, twenty, twenty-one, ...
```

**Actual with `k8v4` + `--kv-resident`** at `--max-context 262144` — 400 tokens, hit the cap, collapsed into a single token:

```
one, two, two, three, three, four, four, five, five, six, six, seven, seven,
seven, seven, seven, seven, seven, seven, seven, seven, seven, ...
```

At `--max-context 32768` (streaming still active) it degenerates the same way, just with a different collapse point:

```
one, two, two, three, three, four, four, five, five, six, six, seven, seven,
eight, eight, nine, eight, nine, nine, ten, ten, eleven, eleven, twelve, twelve,
thirteen, twelve, thirteen, thirteen, ...
```

The 262144 case is reproducible **byte for byte** across repeated runs, so this is deterministic, not sampling noise.

## Isolation matrix

| # | Model | `--kv` | `--kv-resident` | `--max-context` | Needle recall | Short factual (greedy) | 400-token generation |
|---|---|---|---|---|---|---|---|
| 1 | IQ3_XXS | `int8` | 16384 | 262144 | pass | `Red, Blue, Yellow` (6 tok) | clean (156 tok) |
| 2 | IQ3_XXS | `k8v4` | 16384 | 262144 | pass | verbose (32 tok) | **degenerate** (400 tok) |
| 3 | IQ3_XXS | `k8v4` | 0 (off) | 32768 | pass | `Red, Blue, Yellow` (6 tok) | clean (156 tok) |
| 4 | IQ3_XXS | `q4_0` | 16384 | 262144 | pass | `Red, Blue, Yellow` (6 tok) | clean (156 tok) |
| 5 | IQ2_XS (stock) | `int8` | 32768 | 262144 | pass | `Red, Blue, Yellow` (6 tok) | clean (156 tok) |
| 6 | IQ2_XS (stock) | `k8v4` | 32768 | 262144 | pass | verbose (16 tok) | **degenerate** (400 tok) |
| 7 | IQ3_XXS | `k8v4` | 16384 | 32768 | pass | verbose (16 tok) | **degenerate** (400 tok) |

Rows 2 and 6 produce the **same degenerate string**, on two different models (one of them a stock, unmodified quant), so this is not weight-specific.

What the matrix rules out:

- **Not a general "4-bit KV costs quality" effect.** `q4_0` (row 4) quantizes *both* K and V to 4-bit and its output is byte-identical to `int8`; `k8v4` keeps K at INT8 and only rotates V to Q4_0, and it is the one that breaks.
- **Not the k8v4 format itself.** With `--kv-resident 0` (row 3) k8v4 is byte-identical to int8.
- **Not context length.** Row 7 breaks at `--max-context 32768` with streaming on.
- **The streaming path is the variable.** In every broken row `--kv-resident` was active; in every clean k8v4 row it was off. The log confirms streaming was really engaged, e.g. for row 7:

```
strata generate: KV streaming: 20480 of 32768 cells per QSA layer in VRAM, the K/V in 0.30 GiB of pinned RAM
```

and for row 2:

```
strata generate: KV streaming: 20480 of 262144 cells per QSA layer in VRAM, the K/V in 2.39 GiB of pinned RAM
```

## The memory saving itself is real

For what it is worth, k8v4 does deliver the expected memory reduction, which is why I tried it:

| `--kv` (resident 16384, ctx 262144) | K/V in pinned RAM | expert slots | VRAM free, everything loaded |
|---|---|---|---|
| `int8` | 3.09 GiB | 3547 | 524 MiB |
| `k8v4` | **2.39 GiB** (−0.70 GiB, −23%) | 3572 | 558 MiB |

So the format works as advertised for capacity; only the streamed decode path is wrong.

## Why it is easy to miss

- **No error, no warning, no NaN in the engine log.** Output just quietly degrades.
- **Retrieval is unaffected.** My needle test (a 19,357-token prompt with two unique codes at roughly 10% and 90% depth) passed in **all seven** configurations, including the broken ones. The damage only shows up in free generation, and mainly once a reply runs past a few dozen tokens — a short chat answer can look merely "more verbose" (see the "Short factual" column) rather than obviously wrong.

## Workaround

Keep `--kv int8` (or `q4_0`) when `--kv-resident` is in use. `k8v4` appears to be usable only with streaming off, which at large contexts is not what it is wanted for.

Since the 0.1.40 notes say setup now offers k8v4 together with `--kv-resident`, users who take that option will get a silently degraded model.

## Notes

- I could not run `tools/kv_precision_compare.py` for a teacher-forced KL number: the 0.1.40 source tree does not contain that file (the bench README under `bench/results/2026-09-27-kv-q4/` references it). A teacher-forced comparison of k8v4 + `--kv-resident` against int8 would show the divergence properly, if it helps.
- These probes are a degeneration check, not a precision benchmark. They separate "grossly broken" from "not grossly broken"; they are **not** sensitive enough to resolve the subtle perplexity differences the repo's own bench reports for `q4_0`. "q4_0 output identical to int8 here" does not mean q4_0 is lossless.
- Happy to run any specific test or build if useful — the machine is available.

関連リンク

インストール・モデル・リリースへの站内リンク。