Pull requests / #111
#111 Add secondary GPU expert store (--vram-experts) for low-RAM workstatins but with multi-GPU
closed · @lukmanfauzie · 0 Kommentare · Auf GitHub
BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsWindows
Beschreibung
On workstations with 32 GB of system RAM running Qwen3.8-Flash-Next, the full expert set cannot be cached in RAM, causing continuous and costly SSD page reads during prompt prefill. This PR introduces a secondary-GPU VRAM tier (--vram-experts) that stores whole layers of experts on an idle second card (e.g., dual RTX 4060 Ti 16GB setups) and streams them via high-bandwidth PCIe D2H during prefill: 1. SECONDARY-GPU EXPERT STORE (--vram-experts): - VramExpertStore: Allocates resident layers on a second device (default device 1) greedily ordered by routing frequency until VRAM is filled. - Serves high-speed D2H transfers into pinned staging buffers during prompt evaluation (~4.1 GB/s PCIe vs ~220 MB/s sequential disk mmap faults). - Complementary placement: Layers held on the secondary GPU stay in GPU 0's decode cache, avoiding CPU fallback regressions during decode. 2. PREFILL STREAMING & CHUNKING INTEGRATION: - Wired vram_held and vram_fetch into the 384-slot prefill streaming ring (stream_all / issue_until), preventing prompt chunks from bypassing the store. - Supported with --prefill 8192 chunking, reducing prompt expert re-read volume by up to 8x. 3. SETUP, PRUNED MODEL SUPPORT & TELEMETRY: - Configurable via --vram-experts, --vram-expert-layers N, and --vram-expert-device D. - setup.py: Integrates --vram-experts into the --low-ram configuration path. - Supports 256-of-512 pruned expert variants (GSQ-RCO Coder) and custom expert profiles via tools/make_expert_profile.py. - Added prefill telemetry reporting: batched computation time, staging time, and store blob hit percentage in the console and web monitor. ### Empirical Results & Direct Comparison Metric / Feature **Our Dual-GPU Build (`--vram-experts`)** **Upstream v0.1.21 (`--layer-split`)** **Prefill Speed** (22,278 tokens) **524.0 tok/s** (42.5 s prefill time) **Cannot run / Fails startup** **Time to First Token (TTFT)** **45.7 s** N/A (Failed) **Decode Speed** (16 tokens) **17.4 tok/s** (919.9 ms) N/A (Failed) **Expert Cache Hit Rate** (GPU 0) **59.6%** (7,146 / 12,000 requests) 0.0% **Secondary GPU Usage** **GPU 1 holds 22 whole layers (14.5 GB)** in VRAM for high-speed PCIe D2H prompt streaming Aims for pipeline split (layers 0..K-1 on GPU 0, K..47 on GPU 1) **Low-RAM Compatibility** **Full support**: Designed for 32 GB RAM with `--mmap-experts` SSD paging **Incompatible**: Explicitly rejects `--mmap-experts` **System RAM Footprint** Fits comfortably within 32 GB physical RAM Demands 24–34 GB RAM resident arena **Windows WDDM Behavior** Cleanly creates CUDA contexts & buffers **`cudaMemGetInfo` reports `out of memory`**; fails to allocate logits buffer
Mehr auf der Site
Links zu Install, Modellen, Releases.