Pull requests / #1011
#1011 kv: one pinned KV pool shared by the sessions (--kv-pool-tokens; 4 lanes x 256K pin 6.24 GiB instead of 15.5)
closed · @xdbxdbx · 0 comentários · No GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quantsWindows
Descrição
## Problem With KV streaming (`--kv-resident`) every session pins the whole context's K/V in RAM: the main session and each `--batch` slot. At `--max-context 262144` that is 3.09 GiB per session, so 4 lanes pin 5 x 3.09 = 15.5 GiB, although the conversations together rarely fill one context, let alone five. ## Change `--kv-pool-tokens N` (with `--kv-resident`): the sessions share one pinned pool of N cells for the streamed QSA layers. A session holds only the chunks (4096 cells) its conversation has reached. - **Chunk tables:** each session has a chunk table (device copy at a fixed address, so captured graphs keep working, plus a host mirror). Every host-addressing site translates block -> pool block: the decode kernels (`qsa.cu`, `kv_q8.cu`, `kv_q4.cu`), the stream movers (`kv_stream.cu`), prefill, and conversation snapshots, which split copies at chunk boundaries. Unreserved entries point at a trash chunk, so a stray write never lands in another session's cells. - **Moves, not copies:** a conversation moves between the main session and a slot (admission, and a slot handing its conversation back) by swapping chunk tables. A prompt read that gives way (BYIELD) still copies. - **When the pool is full:** - an idle slot's cached conversation gives way, oldest first; - a new request is refused: `ERR KV pool full`, answered as **503 with Retry-After** (`code: kv_pool_full`; in a stream, the error event carries the code, and Anthropic streams use `overloaded_error`); - a decoding lane that can't grow ends with finish `pool`, reported as `"truncated": true` (llama.cpp's field). - **RoPE:** the slot sessions borrow the main session's RoPE table (64 MiB per slot at 262K). - **Metrics:** `/metrics` reports the pool's use (`kv_pool`: cells, free, held per session). - **Server fix:** an `ERR` for a request now ends it at once. The server used to wait 300 s for a `DONE`/`BADM` that never comes after a refusal. Without `--kv-pool-tokens` there are no tables, and every site uses the identity layout as before. ## Validation RTX 4090, 64 GB RAM, Qwen3.8-Flash-Next IQ2_XS, `--resident-experts --kv int8 --kv-resident 32768`; single runs. - **Pinned K/V:** the 524288-cell pool pins 6.24 GiB, against 15.5 GiB for 4 lanes + main without it. - **4 lanes x 256K, 512K pool, more than the pool holds:** 4 prompts of 140K-151K tokens (586K in all) all answered 200. When the pool ran short, a finished lane's cached conversation gave way. - **Admission by move:** 52-79 ms for 140K-151K tokens, against 1.0-1.1 s copied (133K: 1128 ms, 143K: 1020 ms). - **Pressure** (64K context, 128K pool, 4 lanes x ~40K prompt, up to 1500 tokens each): two lanes ran to 1500 tokens; one ended at 45049 cells with `truncated: true`; the fourth request, arriving while three lanes decoded, got 503 `kv_pool_full` with `Retry-After: 10` as soon as its turn came. - **Throughput with the pool**, 4 lanes x 256K: 83.0 tok/s total on code, 81.5 on prose. - **RAM:** available RAM stayed at 15.0 GB or more, with no page-file writes. ## Note At 4 x 256K with resident experts, a 64 GB Windows PC runs close to its commit limit (RAM + page file). With an automatically sized page file (10 GB here, a 73.9 GB limit), the 25 GB resident expert copy was sometimes refused at start. A fixed 32 GB page file removes that; nothing is written to it.
No site
Links install, modelos, releases.