Pull requests / #111

#111 Add secondary GPU expert store (--vram-experts) for low-RAM workstatins but with multi-GPU

closed · @lukmanfauzie · 0 commentaires · Sur GitHub

BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsWindows

Description

On workstations with 32 GB of system RAM running Qwen3.8-Flash-Next, the full expert set cannot be cached in RAM, causing continuous and costly SSD page reads during prompt prefill.

This PR introduces a secondary-GPU VRAM tier (--vram-experts) that stores whole layers of experts on an idle second card (e.g., dual RTX 4060 Ti 16GB setups) and streams them via high-bandwidth PCIe D2H during prefill:

1. SECONDARY-GPU EXPERT STORE (--vram-experts):
   - VramExpertStore: Allocates resident layers on a second device (default device 1) greedily ordered by routing frequency until VRAM is filled.
   - Serves high-speed D2H transfers into pinned staging buffers during prompt evaluation (~4.1 GB/s PCIe vs ~220 MB/s sequential disk mmap faults).
   - Complementary placement: Layers held on the secondary GPU stay in GPU 0's decode cache, avoiding CPU fallback regressions during decode.

2. PREFILL STREAMING & CHUNKING INTEGRATION:
   - Wired vram_held and vram_fetch into the 384-slot prefill streaming ring (stream_all / issue_until), preventing prompt chunks from bypassing the store.
   - Supported with --prefill 8192 chunking, reducing prompt expert re-read volume by up to 8x.

3. SETUP, PRUNED MODEL SUPPORT & TELEMETRY:
   - Configurable via --vram-experts, --vram-expert-layers N, and --vram-expert-device D.
   - setup.py: Integrates --vram-experts into the --low-ram configuration path.
   - Supports 256-of-512 pruned expert variants (GSQ-RCO Coder) and custom expert profiles via tools/make_expert_profile.py.
   - Added prefill telemetry reporting: batched computation time, staging time, and store blob hit percentage in the console and web monitor.

### Empirical Results & Direct Comparison

Metric / Feature
**Our Dual-GPU Build (`--vram-experts`)**
**Upstream v0.1.21 (`--layer-split`)**

**Prefill Speed** (22,278 tokens)
**524.0 tok/s** (42.5 s prefill time)
**Cannot run / Fails startup**

**Time to First Token (TTFT)**
**45.7 s**
N/A (Failed)

**Decode Speed** (16 tokens)
**17.4 tok/s** (919.9 ms)
N/A (Failed)

**Expert Cache Hit Rate** (GPU 0)
**59.6%** (7,146 / 12,000 requests)
0.0%

**Secondary GPU Usage**
**GPU 1 holds 22 whole layers (14.5 GB)** in VRAM for high-speed PCIe D2H prompt streaming Aims for pipeline split (layers 0..K-1 on GPU 0, K..47 on GPU 1)

**Low-RAM Compatibility**
**Full support**: Designed for 32 GB RAM with `--mmap-experts` SSD paging **Incompatible**: Explicitly rejects `--mmap-experts`

**System RAM Footprint**
Fits comfortably within 32 GB physical RAM
Demands 24–34 GB RAM resident arena

**Windows WDDM Behavior**
Cleanly creates CUDA contexts & buffers
**`cudaMemGetInfo` reports `out of memory`**; fails to allocate logits buffer

Sur le site

Liens install, modèles, releases.