Pull requests / #1614
#1614 One shared KV pool for all lanes (`--kv-pool-tokens`): two 256k-context lanes in one 512k-token pool on two 20 GB cards, KV memory follows the tokens in use
open · @noon-at-cgn · 0 comentarios · En GitHub
BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentationWindows
Descripción
## Title
Issue: Related: #1011 (xdbxdbx, the original; closed by GitHub when `main` was force-pushed on 2026-10-06, not rejected, not refiled). Related, different mechanism: #1232 / #1231 (`--kv-grow` with `--batch`, single GPU), #1269 (restore a session file without its K/V in RAM; touches the same snapshot file).
## Summary
**Why it matters:** with KV streaming (`--kv-resident`) every session pins the K/V of its whole `--max-context` in RAM: the main session and each `--batch` slot, however short their conversations are. With `--kv-pool-tokens N` the sessions share one pinned pool of N cells and hold only the 4,096-cell chunks their conversation has reached. On two 20 GB cards that is what lets UD-Q4_K_XL serve **two lanes of 256k context from one 512k-token pool**, 6.24 GiB pinned, with KV memory proportional to the tokens in use instead of lanes x max-context. Without the flag nothing changes: no tables, every site uses the identity layout.
| Measured (2x RTX 3080 20 GB, layer split, UD-Q4_K_XL, `--kv int8 --kv-resident 32768 --kv-pool-tokens 524288`, engine 0.1.40.3 + this change) | |
|---|---|
| Pinned for the pool | 6.24 GiB for 524,288 cells (both stages together) |
| Two lanes (`parallel 2`, `--max-context 262144`) | share the one pool; this is our production config, in use for days |
| Four lanes (`parallel 4`), four concurrent 133,865-token prompts (556k tokens against the 524k pool, a 233k conversation already held) | all answered 200, every lane returned its own needle; the "KV pool full: slot N gives back its cached conversation" path ran |
| A 134k-token conversation moved between the main session and a slot | about 0.2 s (chunk-table swap, no copy) |
From #1011 (xdbxdbx; RTX 4090, 64 GB, IQ2_XS, single runs): 524,288 cells pin 6.24 GiB; admission by move 52-79 ms for 140-151k tokens against 1.0-1.1 s copied; 4 lanes x 256k at 83.0 tok/s total on code; with a small pool, a full pool gave 503 `kv_pool_full` with `Retry-After: 10` and one lane ending `truncated`. Without a pool, the per-session pinning is arithmetic, not a run: (slots + 1) x 3.09 GiB at 262,144 cells, 15.5 GiB for four slots.
## What changed
- Pool: `kv_pool.hpp/.cpp` (new) pin N cells per QSA layer once; each session has a chunk table (device copy at a fixed address so captured graphs keep working, plus a host mirror) and every host-addressing site translates block to pool block: `qsa.cu`, `kv_q8.cu`, `kv_q4.cu`, `kv_stream.cu`, the prefill kernels, `layer.cpp`, `session.cpp`, `mtp.cpp`. Unreserved entries point at a trash chunk, so a stray write never lands in another session's cells.
- Moves, not copies: a conversation moves between the main session and a slot (admission, and a slot handing it back) by swapping chunk tables (`generate.cpp`). A prompt read that gives way (`BYIELD`) still copies.
- When the pool is full: an idle slot's cached conversation gives its K/V back (order in `kv_pool_policy.hpp`, a tested function); then a new request is refused (`ERR KV pool full`, HTTP 503 with `Retry-After`, `kv_pool_full`; `overloaded_error` on `/v1/messages`); a decoding lane that cannot grow ends `length` with `"truncated": true`. `serve/server.py` also reports the pool in `/metrics` (`live.kv_pool`).
- Added after #1011 (ours): the pool across a layer split (`kv_chunks.hpp/.cpp`: one chunk table per conversation, a device copy per stage; #1011 refused a split); K8V4 halves and the session file's K/V reader carry the chunk table, and a session RESTORE reserves its chunks first (`conversation_snapshot.cpp`); pipelined windows back the whole context first and slot sessions borrow RoPE only from a stage that has QSA layers; `kv_stream.hpp` uses `std::size_t`.
- A fix the pool needs (`conversation_cache.hpp`, `generate.cpp`): after an admission move, main's session was cleared but the K/V reuse image retained at a cache restore stayed, so the next park of another conversation could reuse its bytes (silent wrong K/V in a parked entry). The reuse image now carries the tokens and image keys of its conversation and is handed back only for the prefix both share. Test in `conversation_cache_test`.
- Tests: `tests/core/kv_chunks_test.cpp` (CPU, new), `serve/test_kv_pool.py` (new, 16 tests: 503 paths, `truncated`, the POOL line), `conversation_cache_test` additions. Docs: `docs/BATCHING.md` (section and the measurements above), `docs/DETAILS.md`, `docs/MULTI_GPU.md`.
- Left out because main already has them: the `ERR` ends a request at once commit of #1011 (now `ef6c18a`).
## Extra Notes
Credit: the first four commits (the pool, the chunk-swap moves, two comments, the review fixes) are xdbxdbx's from #1011 and keep their author. #1011 was closed automatically by the history rewrite; this is the same work rebased onto 0.1.41 with the review fixes ("pool review fixes" commit) and the additions above. Their branch is `kv-shared-pool`.
Machine: 2x RTX 3080 20 GB (PCIe 3.0, no NVLink), Xeon E5-2696 v4, ~100 GiB usable pinned RAM. The numbers above were taken on 0.1.40.3 plus the pool (our production binary); this branch is cut from fb58e0d and was built and tested on its own on 0.1.41 (CPU tests and serve tests only, no engine was started on it).
Not tested: HIP, SYCL (`sycl/` has its own copies of these headers and is not touched), three or more cards, `k8v4` with the pool at runtime, parking and session files with the pool on these cards (CPU tests only), a single 512k conversation (needs YaRN). **No comparison against per-lane buffers was run**: decode and prompt speed with and without the pool on the same box is the evidence still to collect. Greedy output between runs is not reproducible on our config (the adaptive expert cache), so a solo-vs-slot exactness check of the pool was not possible; cross-lane isolation was checked with distinct needles.
En el sitio
Enlaces a install, modelos, releases.