Pull requests / #1040

#1040 Elastic K/V (--kv-grow, opt-in): the K/V takes VRAM as the context grows, the expert cache holds the rest (+68% cache on a 16 GB card at 262K)

closed · @sergqwer · 0 comments · View on GitHub

BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quantsWindows

Description

Replaces #378. GitHub closed it on 2026-10-05, when my fork was made private by mistake: that took the fork out of the network for good, so it can no longer open pull requests. This is the same branch and commit (`d5ee0cc`), opened from a new fork; the discussion and the measurements are in #378.

---

At `--max-context 262144` the K/V of the whole context is allocated at start: 3.35 GiB at int8 with the drafter's. The expert cache is sized around it. On a 16 GB card that is a quarter of the VRAM, and most requests never reach those cells.

With `--kv-grow` (or `STRATA_KV_GROW=1`), the K/V holds physical memory only for the cells the requests reach. The expert cache holds the rest. **It is off by default:** more slots move experts between the GPU and the CPU, which changes tokens, so it does not pass a byte-identical gate against main by design.

How it works:
- **Virtual memory (`core/vmm.hpp`).** The K/V pools and the cache arena are CUDA VMM ranges: addresses reserved once, physical memory mapped in 2 MiB chunks. The driver calls go through `cudaGetDriverEntryPointByVersion`, so nothing links `cuda.lib`.
  - Nothing moves, so no graph is captured again.
  - `--max-context` cells always fit: the window is never smaller.
- **Growing.** The K/V maps the chunks for 16384 cells at start and grows in steps of 8192.
  - It takes chunks from the slots just below the prompt path's loan.
  - A hotter expert in such a slot first moves (one D2D copy) into the slot of the coldest resident in the loan region, so the cache gives up its coldest experts, as a smaller cache would.
  - The CPU computes the experts given up, like any miss.
- **Giving back.** A short request after a long one hands the chunks back once the outgoing session is parked. The slots refill from the profile.
- **Where it grows:**
  - serve: at a request (its prompt, and a parked conversation before its restore) and in the decode rounds when a window would pass the mapped cells;
  - generate: once, for the prompt plus `--max-new`.
- **When it applies:** one GPU, a profile, the whole K/V in VRAM (no `--kv-resident`), every expert in RAM (not the low-RAM tier). Otherwise the flag is ignored.
- **Test switches:** `STRATA_KV_GROW_INIT` / `_STEP` / `_FLOOR`. `tests/core/vmm_test` maps, moves a chunk between ranges keeping its bytes, and releases.

One pitfall worth knowing: a device-to-device `cudaMemcpy` does not wait for the copy. The moves sync before their source chunks are unmapped. Without that, the copy faulted, and the error surfaced at the next `cudaMemset` as "no VRAM".

## Same output when nothing moves

Setup: IQ2_XS (ISTA), RTX 5090, a 32K prompt and 64 tokens, fixed `--expert-cache 14900`. First-token logits were dumped with #276's `STRATA_DUMP_FIRST_LOGITS` in a test build.

These are bit-identical to 0.1.31, logits and tokens:
- without the flag;
- `--kv-grow` mapped whole (`STRATA_KV_GROW_INIT=262144`);
- `--kv-grow` grown from 4096 cells into new VRAM (`_FLOOR` above the cache).

Grown by giving up 266 slots (171 experts moved): the same first token and the same 64 tokens here. In general a given-up expert is computed on the CPU, which rounds differently.

## Speed

Setup: IQ2_XS (ISTA), RTX 5090, 9950X3D, 128 GB, Windows 11, `--max-context 262144 --kv int8 --prefill auto --spec 4`, cache auto, 256 tokens. The runs alternated with and without the flag.

| | 0.1.31 | `--kv-grow` |
| --- | ---: | ---: |
| expert cache | 15,119-15,207 slots | 17,300-17,503 |
| short chat | 15.06-15.33 ms a round, 161-163 tok/s | **14.71-14.83 ms, 164-167 tok/s** |
| **an emulated 16 GB card** (`--vram-reserve-mib 17884`): expert cache | 3,396 slots | **5,719** |
| 16 GB: short chat | 24.1-25.5 ms a round, 100 tok/s | **21.1 ms, 115 tok/s** |
| 16 GB: a 32K prompt | 7.44 s | **6.36 s** |
| 16 GB: decode after it | 24.6 ms a round | **22.2 ms** |

On our NVFP4 fork (2.76 MB experts, 5090) the cache grows from ~7,060 to ~8,290 slots and a short chat from 111-115 to 126-133 tok/s.

## Serve

A served run with `--conversation-cache-mib 8192` behaved as expected:
- a chat;
- a 17.8K prompt and 4,000 tokens: grown 16K -> 20K at the request and 24K mid-decode;
- a 31.7K document: grown to 32K;
- two short chats: trimmed to 4K, 159 slots refilled;
- the document's follow-up: grown before its parked conversation was restored (31,749 tokens reused).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.