Pull requests / #265
#265 nvme kv cache v2 - the delta tier + v4 snapshots (replaces #52)
closed · @maedoc · 0 Kommentare · Auf GitHub
Beschreibung
Builds on the work reviewed in #52, with the two asks it left open addressed: snapshots bound to the weights that wrote them, and `--kv k8v4` in the format key. ## Summary The v3 header carried the model geometry, and geometry is not identity - two weight sets of one architecture share it exactly, so a snapshot of another model's conversation had everything it needed to match, verify its digest and apply cleanly. v4 makes the header say whose bytes these are. - **v4 snapshot header** - carries the weight-set fingerprint (§5.8) and the KV storage key, so a same-geometry foreign-weights snapshot can never promote; both refusals are the recoverable class, before the first CUDA call, so a mismatch re-prefills and never mixes state - **`--kv k8v4` in the format key** - the key is now `qsa_kv_key`, total over all four formats, so a k8v4 store is keyed apart from every other layout instead of being filed as fp16 - **delta tier** - turn-boundary incremental writes (content-addressed chunks plus a per-turn state record), so the write cost tracks the new tokens rather than the session; off unless `--kv-nvme` is passed - rebased onto 0.1.28 Two things worth flagging: this converges onto #189's shared core as soon as that lands - this branch currently carries the imported copy - and the disk tier in #190 is acknowledged; the coordination is the open item. The one re-prefill per conversation after the upgrade is the deliberate cost of the version bump: a v3 file is refused by version, never converted.
Mehr auf der Site
Links zu Install, Modellen, Releases.