Pull requests / #37

#37 core: elastic expert cache via CUDA VMM and on-demand GPU vision lifecycle

closed · @code-martin · 0 コメント · GitHub で見る

NVIDIA / CUDAModels & quantsWindows

本文

### Summary
This PR implements elastic resizing for the MoE expert cache using the CUDA Virtual Memory Management (VMM) API and introduces an on-demand lifecycle for GPU Vision models.

### Key Changes
- **CUDA VMM Elastic Expert Cache**:
  - Replaces fixed monolithic cudaMalloc allocations with cuMemAddressReserve, cuMemCreate, and cuMemMap / cuMemSetAccess.
  - Supports dynamic online shrink (RESIZE) and expansion (GROW) without reallocating virtual addresses or invalidating captured CUDA graphs.
  - Automatically falls back to standard cudaMalloc if VMM is unsupported on older architectures.
- **On-Demand GPU Vision Lifecycle**:
  - Vision encoder is dynamically loaded into VRAM only when image embeddings are required.
  - Temporarily shrinks the expert cache to release VRAM, performs vision encoding, offloads the vision model, and instantly grows the expert cache back to full capacity.
  - Reclaims over 600 MoE slots (~830 MB) of VRAM permanently during text-only generation.
- **Client Validation Tool**:
  - Adds 	ools/test_vision_client.py for end-to-end verification of image encoding, VRAM contraction, and cache restoration.

### Testing
- Verified on Windows with MSVC C++20 and CUDA 12.6 (RTX 3080 10GB).
- Text-only generation runs with 2,459 expert slots (up from 1,840). Vision requests encode in ~7.5s and immediately restore the full 2,459 cache slots.

関連リンク

インストール・モデル・リリースへの站内リンク。