Pull requests / #1232

#1232 enable --kv-grow with --batch

open · @BlueKingMuch · 0 Kommentare · Auf GitHub

NVIDIA / CUDAModels & quantsWindows

Beschreibung

Builds on #1231 (the first commit); this PR is the second.

With `--batch N` every slot carves a session for the whole `--max-context`, used or not: 3.65 GiB each at 262K with int8 K/V, or 0.95 GiB with `--kv-resident 32768` plus ~3.1 GiB of pinned RAM. 

`--kv-grow` switches itself off beside `--batch` ("the batch slots carve their own K/V"). This makes every slot session elastic under `--kv-grow`.

## What it does

- The elastic pools come in groups: the main session and its drafter are group 0, batch slot b's session group b + 1 (`qsa_set_kv_elastic_group` while it is initialized). `qsa_kv_elastic_{cells,need,grow,shrink}` take the group; every existing caller is group 0, so one session runs exactly as before.
- With `--batch-mtp` each slot's drafter K/V joins its slot's group, so it grows and shrinks with the slot.
- `kvg_ensure` / `kvg_trim` take the group too: all groups draw on the same loan of the expert cache (the slots just below the prompt path's) and give their chunks back to it.
- A slot grows before every write: at its admission (`copy_to_slot`, trimming what it held beyond first), at a BYIELD, and before each batch window (its position + 1 + 64). A slot starts with 4,096 cells mapped (`STRATA_KV_GROW_SLOT_INIT`).
- A slot's indexer keys (the pooled rows) are an array of its elastic range too, one row per page. The main session's stay whole: the prompt path rounds all of them for its block scores; the windows read rows only up to the current block, and the snapshot code copies only the rows below `upto`.
- `--kv-grow` stays off beside `--vram-elastic` and `--peer-device`, as before. Without `--kv-grow` nothing changes.

## Tests

- Every slot token for token like alone, `tools/batch_test.py`, four slots, `--pcie-frac 0 --adapt-every 1000000
  --no-prefill-borrow --prefill 1024`: without and with `--kv-grow` 4 of 4 IDENTICAL; with `--batch-mtp`, without and with `--kv-grow`, 4 of 4 IDENTICAL (with the one-line fix from #1063: the MTP directory here has a `draft_vocab.bin`).
- `tools/batch_interleave_test.py` with `--kv-grow` and the slots pre-grown (`STRATA_KV_GROW_SLOT_INIT=16384`): 8 of 8 IDENTICAL.
- `vmm_test`, `conversation_snapshot_test`, `conversation_cache_test`, `conversation_file_test` pass.

## Measured

On 0.1.40.1, this against 0.1.40.1: RTX 4080 SUPER 32 GB, PCIe 3.0, Windows 11, IQ3_S, 262K int8, 4 slots,
conversation cache 12,288 MiB; 1 orchestrator + 4 subagents, two rounds, the orchestrator working on its own between them (6 turns a round):

| | whole run | experts in VRAM | a slot session | RAM free after start |
|---|---|---|---|---|
| 0.1.40.1, `--kv-resident 32768` | 375 s | 9,958 | 0.95 GiB + 3.09 GiB pinned | 20.0 GB |
| this PR, `--kv-grow` | 358 s | 11,402 | 0.18 GiB | 37.4 GB |

Each subagent's slot grew to 16,384 cells and took 86-88 expert slots for it; a slot that held the orchestrator grew to 49,152 cells (227-314 slots) and gave 223 back when it went on with a subagent.

The same pair with the VRAM of a 16 GB card, emulated on this card (`--vram-reserve-mib 17084`: the same PCIe link, 16 GiB less for the cache):

| | whole run | experts in VRAM |
|---|---|---|
| 0.1.40.1, `--kv-resident 32768` | 662 s | 2,148 |
| this PR, `--kv-grow` | 629 s | 3,581 |

cc @sergqwer - this extends your elastic K/V (#378) to the batch slots.


Mehr auf der Site

Links zu Install, Modellen, Releases.