Issues / #1169

#1169 k8v4 + --kv-resident: the batched KV append (#783) skips the host mirror for hybrid — parking and KV streaming serve stale bytes (cross-conversation contamination on 0.1.40)

closed · @taxah92 · 1 comments · View on GitHub

Setup & installServer & APINVIDIA / CUDAModels & quants

Description

Field report from a single-GPU production rig, stock v0.1.40 (commit `1cbcacb`), no fork, no layer split. The same code is still on current `main` (`82f46a8`): `src/core/verify.cpp` lines 949-953.

**Setup:** Tesla V100 SXM2 16 GB (sm_70), CUDA 12.4 distro toolkit, Ubuntu 26.04, EPYC 7532. IQ3_S pack, `--max-context 1048576 --rope-scaling yarn --kv k8v4 --kv-resident 32768`, `--mtp --spec 3 --suffix-draft 3`, conversation parking on (`--conversation-cache-mib 24576`, 3 slots), `--prompt-cache 16`. Two long agent conversations (~95k and ~150k tokens) alternating through parking.

**Symptom** (appeared within ~1 hour of switching `--kv int8 → k8v4`, vanished on revert): one conversation's replies started containing material from the *other* conversation — content that was never in its own prompt — while losing its own most recent decode-written turns (the model "forgot" work it had just done and tried to redo it). The int8 era of the same engine, same flags except `--kv`, never showed this.

**Where:** `src/core/verify.cpp`, the batched append branch of the verify window (#783, default on, taken whenever the MTP window has n>1 rows):

```cpp
if (st.kv_hybrid) {            // K8V4
    kv_append_q8_steps(st.k_q, st.k_q, st.k_scale, st.k_scale, st.page_table, step_b, kStepCount,
                       kc_b, kc_b, (int) (NKV * HD), n, s, cs, nullptr);   // ← host mirror = nullptr
    kv_append_q4_steps(st.v_q4, st.v_q4, st.page_table, step_b, kStepCount, n, vc_b, vc_b, s, cs,
                       nullptr);                                          // ← same
} else if (st.kv_q4)   kv_append_q4_steps(..., &st.host);
else if (st.kv_int8)   kv_append_q8_steps(..., &st.host);
```

The per-row branch just below (`else for (int t = tb; t < te; ++t)`) handles hybrid **correctly** — it builds `kv_hybrid_k_half(st.host)` / `kv_hybrid_v_half(st.host)` and passes them under the `mirror = st.host.present()` guard. The batched branch hardcodes `nullptr`.

**Consequence:** with `--kv-resident` (streaming), every cell written by decode never reaches the pinned host copy. Both consumers of that copy then read stale bytes: the KV-streaming tier (block reads that miss VRAM) and conversation parking / prompt-cache restore ("a streamed one reads its host copy", `conversation_snapshot.cpp`). Restore trusts the invariant that every writer mirrors — so reused prefixes resume from another conversation's arena pages and fresh decode turns vanish. This matches exactly the failure signature @fishlikeX described for the fork's batch replay in #711 (jmnargi/Strata-V100#36: "reused prefixes returned other conversations", five-line fix) — the same class of bug, now in upstream's #783 batched append. As @fishlikeX predicted there: "If the batched replay form ever moves upstream, the hybrid mirror has to move with it — that is the one line this port initially missed."

**Suggested fix:** mirror the per-row branch in the batched one:

```cpp
if (st.kv_hybrid) {
    const KvHostPools hk = kv_hybrid_k_half(st.host), hv = kv_hybrid_v_half(st.host);
    const bool mirror = st.host.present();
    kv_append_q8_steps(..., mirror ? &hk : nullptr);
    kv_append_q4_steps(..., mirror ? &hv : nullptr);
}
```

**Workaround in the meantime:** `STRATA_NO_BATCH_KV_STEP=1` (the #783 PR-d kill switch) routes appends through the correct per-row branch; `--kv int8` is unaffected either way.

Related but separate: #1135 (deterministic repetition under k8v4 on sequential generation) — same format, different defect. Both argue for a k8v4 + streaming + parking case in the test suite; `tools/stateless_prompt_probe.py --parked` from the V100 fork reproduces this one.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.