Pull requests / #1040
#1040 Elastic K/V (--kv-grow, opt-in): the K/V takes VRAM as the context grows, the expert cache holds the rest (+68% cache on a 16 GB card at 262K)
closed · @sergqwer · 0 commentaires · Sur GitHub
BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quantsWindows
Description
Replaces #378. GitHub closed it on 2026-10-05, when my fork was made private by mistake: that took the fork out of the network for good, so it can no longer open pull requests. This is the same branch and commit (`d5ee0cc`), opened from a new fork; the discussion and the measurements are in #378. --- At `--max-context 262144` the K/V of the whole context is allocated at start: 3.35 GiB at int8 with the drafter's. The expert cache is sized around it. On a 16 GB card that is a quarter of the VRAM, and most requests never reach those cells. With `--kv-grow` (or `STRATA_KV_GROW=1`), the K/V holds physical memory only for the cells the requests reach. The expert cache holds the rest. **It is off by default:** more slots move experts between the GPU and the CPU, which changes tokens, so it does not pass a byte-identical gate against main by design. How it works: - **Virtual memory (`core/vmm.hpp`).** The K/V pools and the cache arena are CUDA VMM ranges: addresses reserved once, physical memory mapped in 2 MiB chunks. The driver calls go through `cudaGetDriverEntryPointByVersion`, so nothing links `cuda.lib`. - Nothing moves, so no graph is captured again. - `--max-context` cells always fit: the window is never smaller. - **Growing.** The K/V maps the chunks for 16384 cells at start and grows in steps of 8192. - It takes chunks from the slots just below the prompt path's loan. - A hotter expert in such a slot first moves (one D2D copy) into the slot of the coldest resident in the loan region, so the cache gives up its coldest experts, as a smaller cache would. - The CPU computes the experts given up, like any miss. - **Giving back.** A short request after a long one hands the chunks back once the outgoing session is parked. The slots refill from the profile. - **Where it grows:** - serve: at a request (its prompt, and a parked conversation before its restore) and in the decode rounds when a window would pass the mapped cells; - generate: once, for the prompt plus `--max-new`. - **When it applies:** one GPU, a profile, the whole K/V in VRAM (no `--kv-resident`), every expert in RAM (not the low-RAM tier). Otherwise the flag is ignored. - **Test switches:** `STRATA_KV_GROW_INIT` / `_STEP` / `_FLOOR`. `tests/core/vmm_test` maps, moves a chunk between ranges keeping its bytes, and releases. One pitfall worth knowing: a device-to-device `cudaMemcpy` does not wait for the copy. The moves sync before their source chunks are unmapped. Without that, the copy faulted, and the error surfaced at the next `cudaMemset` as "no VRAM". ## Same output when nothing moves Setup: IQ2_XS (ISTA), RTX 5090, a 32K prompt and 64 tokens, fixed `--expert-cache 14900`. First-token logits were dumped with #276's `STRATA_DUMP_FIRST_LOGITS` in a test build. These are bit-identical to 0.1.31, logits and tokens: - without the flag; - `--kv-grow` mapped whole (`STRATA_KV_GROW_INIT=262144`); - `--kv-grow` grown from 4096 cells into new VRAM (`_FLOOR` above the cache). Grown by giving up 266 slots (171 experts moved): the same first token and the same 64 tokens here. In general a given-up expert is computed on the CPU, which rounds differently. ## Speed Setup: IQ2_XS (ISTA), RTX 5090, 9950X3D, 128 GB, Windows 11, `--max-context 262144 --kv int8 --prefill auto --spec 4`, cache auto, 256 tokens. The runs alternated with and without the flag. | | 0.1.31 | `--kv-grow` | | --- | ---: | ---: | | expert cache | 15,119-15,207 slots | 17,300-17,503 | | short chat | 15.06-15.33 ms a round, 161-163 tok/s | **14.71-14.83 ms, 164-167 tok/s** | | **an emulated 16 GB card** (`--vram-reserve-mib 17884`): expert cache | 3,396 slots | **5,719** | | 16 GB: short chat | 24.1-25.5 ms a round, 100 tok/s | **21.1 ms, 115 tok/s** | | 16 GB: a 32K prompt | 7.44 s | **6.36 s** | | 16 GB: decode after it | 24.6 ms a round | **22.2 ms** | On our NVFP4 fork (2.76 MB experts, 5090) the cache grows from ~7,060 to ~8,290 slots and a short chat from 111-115 to 126-133 tok/s. ## Serve A served run with `--conversation-cache-mib 8192` behaved as expected: - a chat; - a 17.8K prompt and 4,000 tokens: grown 16K -> 20K at the request and 24K mid-decode; - a 31.7K document: grown to 32K; - two short chats: trimmed to 4K, 159 slots refilled; - the document's follow-up: grown before its parked conversation was restored (31,749 tokens reused). 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
Sur le site
Liens install, modèles, releases.