Pull requests / #1231

#1231 kv-grow: a run that cannot lend cache slots maps the whole window up front instead of writing past its first 16K cells

open · @BlueKingMuch · 0 comments · View on GitHub

Server & APINVIDIA / CUDAModels & quantsSecurityWindows

Description

With `--kv-grow` the K/V is made elastic at session init and maps its first 16,384 cells (`STRATA_KV_GROW_INIT`).

Growing needs more than that: a device residency table, cache slots and a loan above the floor of 128 slots. 

When a run lacks one of them, `kvg_start` leaves the growth off and nothing ever maps more, so a prompt past 16K writes K/V into unmapped memory.

This is @sergqwer's fix from #378 (d5ee0cc), which the port into 0.1.40 does not have, cherry-picked onto 0.1.40.1 with him as the author:

- a run that cannot lend cache slots maps the whole window at start, from new memory, and says so ("elastic K/V off for this run (no cache slots to lend): the whole window, N cells, mapped up front"), or stops with a message when it does not fit;
- the growth's and the trim's residency-table uploads wait for their own copy (a pageable `cudaMemcpy` may return before its DMA lands, and the streams that read the table are non-blocking).

## Reproduced

On 0.1.40.1 (82f46a8): RTX 4080 SUPER 32 GB, Windows 11, IQ3_S, 262K context, int8 K/V, `--kv-grow
--expert-cache 80` (102 cache slots, below the floor), one conversation of 30,000 tokens and then three turns of +2,300:

| | 0.1.40.1 | this PR |
|---|---|---|
| at start | no elastic K/V line (growth off, 16,384 cells mapped) | the whole window, 262,144 cells, mapped up front |
| the 30,000-token prompt | `prefill copy_i32: an illegal memory access was encountered` | read in 11.4 s |
| the turns after it (to 37,092 tokens) | - | resumed, 2,300-2,369 tokens read each |

The second part (the uploads) did not fail here; it is in the PR because it is part of the same fix in #378.

cc @sergqwer

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.