Issues / #1495
#1495 [Feature]: experimental per-stage KV-grow for a two-GPU layer split at native 262K
open · @k93k2J-glitch · 0 comentarios · En GitHub
Setup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsSecurityWindows
Descripción
## Goal Keep the native 262,144 context capacity on a heterogeneous 16GB + 16GB GPU pair while allowing shorter conversations to use unused KV physical memory for expert residency. Both GPUs participate through a 32/16 layer split with pipeline windows inside one serial conversation. Current main's KV-grow condition excludes `multi_gpu`. This is a request to discuss an experimental extension, not a report that the documented unsupported combination is a stock bug. ## Related work and proposed scope #1223 addresses owning-device VMM allocation for `--peer-device`, while keeping layer splits excluded. Our local prototype overlaps its ownership work but additionally uses per-device QSA pools, fixed drafter-KV accounting, per-stage KV/expert loan controllers and layer-scoped numeric slot ownership. Budget changes refresh every stage's residency table while keeping virtual addresses stable for captured graphs. Elasticity is enabled before target/draft/stage initialization, and the request output budget is mapped before decode. The newer optional KV-grow hold behavior is preserved. It would be useful to coordinate with #1223 rather than submit another competing ownership implementation. #765 is related to budgeting, and #1231 covers small-cache/no-loan safety that this normal two-card setup has not tested. ## Existing evidence Local Release source build based on v0.1.40.3; CUDA 13.0.88 / SM120; RTX 5070 Ti 16GB + RTX 5060 Ti 16GB, no P2P as previously checked on the pair; Ryzen 7 9800X3D 8 cores / 16 threads; Windows WDDM, driver 616.56. Two variants each ran actual 65,536 -> 131,072 -> 261,000 -> 8,192 input tokens, with actual 1,024 generated tokens each and no prefix reuse. All eight outputs passed four functional cases on one fixed Python task. On both variants, allocated KV changed 2,112 -> 160 MiB on GPU0 and 1,320 -> 100 MiB on GPU1 after the long-to-short transition, with resident experts increasing on each card. These are registered KV allocation counters, not total VRAM readings. The complete local prototype also passed standalone VMM ownership, foreign-chunk rejection, stable-address/data-retention checks and dual-GPU DMA/reuse checks. The extracted public prototype patches have passed clean-tag application and core-source/parser equivalence checks, but were not independently rebuilt or run during publication preparation. Companion results-only report: #1494. It includes the synthetic cases, numeric evidence, opt-in portable runner and prototype patches, with local auth, Windows launchers and personal data excluded. ## Boundaries - This prototype is serial serving with pipeline windows, not split batch-MTP support. - Resident-RAM/helper/peer/vision combinations, other splits, other quantizations and other platforms are not validated here. - The older ten-round custom-fork DMA comparison does not prove that this combination is globally optimal or that the new v0.1.40.3 prototype is faster than stock. - Long-distance recall/perplexity have not been tested. The coding workload produces comments after completing the function. - A small atomic warning-guard suggestion is included as a separate artifact. It is a static concern in a concurrent error path, with no observed model failure attributed to it. - A cached-alias query after freeing a test allocation does not demonstrate an upstream use-after-free under the helper's fixed-blob caller assumptions. Temporary-buffer invalidation/reuse protection is described as an extension, not a confirmed stock corruption bug. Would a narrowly scoped experimental layer-split KV-grow extension be useful, after aligning ownership with #1223? If so, the core ownership/QSA/stage-budget changes and regressions can be separated from DMA staging and diagnostics for review.
En el sitio
Enlaces a install, modelos, releases.