Pull requests / #37
#37 core: elastic expert cache via CUDA VMM and on-demand GPU vision lifecycle
closed · @code-martin · 0 コメント · GitHub で見る
NVIDIA / CUDAModels & quantsWindows
本文
### Summary This PR implements elastic resizing for the MoE expert cache using the CUDA Virtual Memory Management (VMM) API and introduces an on-demand lifecycle for GPU Vision models. ### Key Changes - **CUDA VMM Elastic Expert Cache**: - Replaces fixed monolithic cudaMalloc allocations with cuMemAddressReserve, cuMemCreate, and cuMemMap / cuMemSetAccess. - Supports dynamic online shrink (RESIZE) and expansion (GROW) without reallocating virtual addresses or invalidating captured CUDA graphs. - Automatically falls back to standard cudaMalloc if VMM is unsupported on older architectures. - **On-Demand GPU Vision Lifecycle**: - Vision encoder is dynamically loaded into VRAM only when image embeddings are required. - Temporarily shrinks the expert cache to release VRAM, performs vision encoding, offloads the vision model, and instantly grows the expert cache back to full capacity. - Reclaims over 600 MoE slots (~830 MB) of VRAM permanently during text-only generation. - **Client Validation Tool**: - Adds ools/test_vision_client.py for end-to-end verification of image encoding, VRAM contraction, and cache restoration. ### Testing - Verified on Windows with MSVC C++20 and CUDA 12.6 (RTX 3080 10GB). - Text-only generation runs with 2,459 expert slots (up from 1,840). Vision requests encode in ~7.5s and immediately restore the full 2,459 cache slots.
関連リンク
インストール・モデル・リリースへの站内リンク。