Pull requests / #711

#711 kv: --kv k8v4 streams with --kv-resident

closed · @T-Crypt · 0 Kommentare · Auf GitHub

Setup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

Beschreibung

## What

`--kv k8v4` (INT8 K + rotated Q4_0 V, 816 B/cell) now streams with `--kv-resident N`, like fp16, int8 and q4_0. Both refusals are gone (`generate.cpp` and the backstop in `qsa_state_init`).

## How

The streaming machinery already moves a block as a list of byte runs (`runs_of` in `kv_stream.cu`), so K8V4 is a fourth format with three runs, and everything that moves blocks (resolve, copy, ring restore, prompt-path staging) takes it unchanged.

- `KvFormat` gains `kKvHybrid = 3`, the number conversation snapshots already use for K8V4. Its runs are K codes, K scales and V q4_0, and `kv_block_bytes` knows its size. `qsa_kv_format` returns it instead of exiting.
- Writers: the hybrid appends already fold each half onto one pool (`kv_append_q8_step(k_q, k_q, ...)`, `kv_append_q4_step(v_q4, v_q4, ...)`). Two helpers in `kv_stream.hpp`, `kv_hybrid_k_half` and `kv_hybrid_v_half`, give each half's host copy (and staging pool) in the shape those kernels expect. Both fields of a half point at the same array on purpose: the kernels test the **K**-side pointer for presence (`if (host.k_q4 != nullptr)`) and then pick K or V per lane, so a V half with only `v_q4` set would silently skip the host write. The decode (`layer.cpp`), verify-window (`verify.cpp`) and prompt (`prefill.cpp`) appends pass them.
- `qsa_state_init` allocates the host copy in the three runs, and `take_stage` gives the prompt path's staging pool the K8V4 layout (`pools_of` then reads it as mode 3).
- Conversation snapshots read a streamed K8V4 state's host copy, as they already do for the other formats. A ring is still refused.

`setup.py` is unchanged and still keeps K8V4 unstreamed: a prebuilt engine without this refuses the pair, so turning it on there needs a `MIN_ENGINE` that includes it. I left that for you to time.

## Tests

- `kv_stream_parity` now runs K8V4 too, with the engine's folded appends. Streamed and resident attention are bitwise equal: 306 batches, 702,021 block lookups, 70,449 misses, 0 failures, and the ring restore is identical. The other three formats are unchanged and still pass.
- Negative control: with the V half's K-side pointer left null (the trap above), the test fails from the first batch (`k8v4 pos 130 n_q 8: streamed attention differs`).
- `kv_hybrid_parity`: all pass.

## Measured

RTX 4090 24 GB (sm_89), i9-14900K, 31 GB RAM, Linux, CUDA 13. Qwen3.8-Flash-Next IQ2_XS, `--max-context 262144 --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp <rt> --resident-experts`, one engine start per row. Long prompts are synthetic log lines with one planted line in the middle, and the answer was right in every row. "Decode at depth" is a ~300-token second turn on the cached prompt.

| `--kv` | streamed | GPU expert slots | short decode | 32K: read / decode | 119K: read / decode | 238K: read / decode |
| --- | --- | ---: | ---: | --- | --- | --- |
| k8v4 | no | 11,208 | 117.7 | 3,879 / 92.8 | 3,870 / 114.2 | 3,658 / 111.0 |
| **k8v4** | **`--kv-resident 32768`** | 12,953 | **126.2** | 3,935 / 99.8 | 3,960 / 116.4 | 3,659 / 108.2 |
| q4_0 | `--kv-resident 32768` | 13,033 | 117.8 | 4,090 / 99.9 | 4,045 / 108.2 | 3,796 / 120.1 |

Streamed K8V4 holds 2.39 GiB of KV in pinned RAM, 95.7% of its KV block reads hit VRAM, and its freed VRAM went to 1,745 more expert slots (11,208 -> 12,953). Prompt-cache checkpoints were restored on the streamed K8V4 state (114,688 and 229,376 tokens reused) and decoded normally.

Single runs, so read the decode differences of a few percent as noise, except the short-decode gain, which comes from the larger expert cache.

## Not tested

- HIP: not built here. The change is in shared C++ and in `kv_stream.cu`, so a HIP build should pick it up the same way it does for the other formats.
- Multi-GPU (`--gpus`). A port of this change on 2x V100 with a layer split reports wrong answers on short stateless prompts under k8v4 that int8 does not show (jmnargi/Strata-V100#34). Not reproduced on sm_89 yet; see the thread below.

## Tested by others

- @midhatn, RTX 4070 Laptop 8 GB, Windows, Swift 1.5 IQ3_XXS, 64K context, `--kv k8v4 --kv-resident 20480`, 2 GiB two-slot conversation cache: a 17-request smoke test passed, including a 57,745-token conversation restore after an unrelated request and an image prefix restore. `kv_stream_parity` and `kv_hybrid_parity` pass on that GPU. Built on v0.1.39 with #732, #733 (opt-ins off) and #749.

Mehr auf der Site

Links zu Install, Modellen, Releases.