Pull requests / #1231
#1231 kv-grow: a run that cannot lend cache slots maps the whole window up front instead of writing past its first 16K cells
open · @BlueKingMuch · 0 comentarios · En GitHub
Server & APINVIDIA / CUDAModels & quantsSecurityWindows
Descripción
With `--kv-grow` the K/V is made elastic at session init and maps its first 16,384 cells (`STRATA_KV_GROW_INIT`).
Growing needs more than that: a device residency table, cache slots and a loan above the floor of 128 slots.
When a run lacks one of them, `kvg_start` leaves the growth off and nothing ever maps more, so a prompt past 16K writes K/V into unmapped memory.
This is @sergqwer's fix from #378 (d5ee0cc), which the port into 0.1.40 does not have, cherry-picked onto 0.1.40.1 with him as the author:
- a run that cannot lend cache slots maps the whole window at start, from new memory, and says so ("elastic K/V off for this run (no cache slots to lend): the whole window, N cells, mapped up front"), or stops with a message when it does not fit;
- the growth's and the trim's residency-table uploads wait for their own copy (a pageable `cudaMemcpy` may return before its DMA lands, and the streams that read the table are non-blocking).
## Reproduced
On 0.1.40.1 (82f46a8): RTX 4080 SUPER 32 GB, Windows 11, IQ3_S, 262K context, int8 K/V, `--kv-grow
--expert-cache 80` (102 cache slots, below the floor), one conversation of 30,000 tokens and then three turns of +2,300:
| | 0.1.40.1 | this PR |
|---|---|---|
| at start | no elastic K/V line (growth off, 16,384 cells mapped) | the whole window, 262,144 cells, mapped up front |
| the 30,000-token prompt | `prefill copy_i32: an illegal memory access was encountered` | read in 11.4 s |
| the turns after it (to 37,092 tokens) | - | resumed, 2,300-2,369 tokens read each |
The second part (the uploads) did not fail here; it is in the PR because it is part of the same fix in #378.
cc @sergqwerEn el sitio
Enlaces a install, modelos, releases.