Issues / #1169
#1169 k8v4 + --kv-resident: the batched KV append (#783) skips the host mirror for hybrid — parking and KV streaming serve stale bytes (cross-conversation contamination on 0.1.40)
closed · @taxah92 · 1 commentaires · Sur GitHub
Setup & installServer & APINVIDIA / CUDAModels & quants
Description
Field report from a single-GPU production rig, stock v0.1.40 (commit `1cbcacb`), no fork, no layer split. The same code is still on current `main` (`82f46a8`): `src/core/verify.cpp` lines 949-953.
**Setup:** Tesla V100 SXM2 16 GB (sm_70), CUDA 12.4 distro toolkit, Ubuntu 26.04, EPYC 7532. IQ3_S pack, `--max-context 1048576 --rope-scaling yarn --kv k8v4 --kv-resident 32768`, `--mtp --spec 3 --suffix-draft 3`, conversation parking on (`--conversation-cache-mib 24576`, 3 slots), `--prompt-cache 16`. Two long agent conversations (~95k and ~150k tokens) alternating through parking.
**Symptom** (appeared within ~1 hour of switching `--kv int8 → k8v4`, vanished on revert): one conversation's replies started containing material from the *other* conversation — content that was never in its own prompt — while losing its own most recent decode-written turns (the model "forgot" work it had just done and tried to redo it). The int8 era of the same engine, same flags except `--kv`, never showed this.
**Where:** `src/core/verify.cpp`, the batched append branch of the verify window (#783, default on, taken whenever the MTP window has n>1 rows):
```cpp
if (st.kv_hybrid) { // K8V4
kv_append_q8_steps(st.k_q, st.k_q, st.k_scale, st.k_scale, st.page_table, step_b, kStepCount,
kc_b, kc_b, (int) (NKV * HD), n, s, cs, nullptr); // ← host mirror = nullptr
kv_append_q4_steps(st.v_q4, st.v_q4, st.page_table, step_b, kStepCount, n, vc_b, vc_b, s, cs,
nullptr); // ← same
} else if (st.kv_q4) kv_append_q4_steps(..., &st.host);
else if (st.kv_int8) kv_append_q8_steps(..., &st.host);
```
The per-row branch just below (`else for (int t = tb; t < te; ++t)`) handles hybrid **correctly** — it builds `kv_hybrid_k_half(st.host)` / `kv_hybrid_v_half(st.host)` and passes them under the `mirror = st.host.present()` guard. The batched branch hardcodes `nullptr`.
**Consequence:** with `--kv-resident` (streaming), every cell written by decode never reaches the pinned host copy. Both consumers of that copy then read stale bytes: the KV-streaming tier (block reads that miss VRAM) and conversation parking / prompt-cache restore ("a streamed one reads its host copy", `conversation_snapshot.cpp`). Restore trusts the invariant that every writer mirrors — so reused prefixes resume from another conversation's arena pages and fresh decode turns vanish. This matches exactly the failure signature @fishlikeX described for the fork's batch replay in #711 (jmnargi/Strata-V100#36: "reused prefixes returned other conversations", five-line fix) — the same class of bug, now in upstream's #783 batched append. As @fishlikeX predicted there: "If the batched replay form ever moves upstream, the hybrid mirror has to move with it — that is the one line this port initially missed."
**Suggested fix:** mirror the per-row branch in the batched one:
```cpp
if (st.kv_hybrid) {
const KvHostPools hk = kv_hybrid_k_half(st.host), hv = kv_hybrid_v_half(st.host);
const bool mirror = st.host.present();
kv_append_q8_steps(..., mirror ? &hk : nullptr);
kv_append_q4_steps(..., mirror ? &hv : nullptr);
}
```
**Workaround in the meantime:** `STRATA_NO_BATCH_KV_STEP=1` (the #783 PR-d kill switch) routes appends through the correct per-row branch; `--kv int8` is unaffected either way.
Related but separate: #1135 (deterministic repetition under k8v4 on sequential generation) — same format, different defect. Both argue for a k8v4 + streaming + parking case in the test suite; `tools/stateless_prompt_probe.py --parked` from the V100 fork reproduces this one.Sur le site
Liens install, modèles, releases.