Pull requests / #151

#151 Add secondary GPU expert store (--vram-experts) and low-RAM multi-GPU setup

closed · @lukmanfauzie · 0 comments · View on GitHub

Setup & installMulti-GPUAMD / HIPModels & quantsWindows

Description

Replaces and supersedes #111. Rebased into a single clean commit on top of latest `main` (`v0.1.24`), addressing all review feedback:

1. **Native router dynamic expert count + batched multi-token launch**:
   - Upstream v0.1.24's batched `native_router_top10_multi` now takes `n_expert` row stride dynamically.
   - Seamlessly handles both standard 512-expert models and 256-expert pruned models (Coder).

2. **MTP drafter**:
   - Preserves upstream v0.1.24's bulk `FILE*` loader in `mtp.cpp` while taking probed draft expert count.

3. **Native pack `--mmap-experts` support**:
   - Native IQ packs shipping `experts.bin` (via `tools/iq_pack.py --experts-bin`) map directly. `FileExpertSource::offset_of` indexes blobs using `ExpertLayout::blob_offset`, addressing both uniform and per-layer variable layouts.

4. **Server tests & telemetry preserved**:
   - Restored all 8 server tests; the entire test suite (38/38 server tests, 73/73 unit/integration tests) passes cleanly.
   - Live prompt throughput indicators (`prompt_tok_s`, `prompt_eta_s`) feed to `app.js` via upstream's `PP` pipeline.

5. **Dual-GPU Low-RAM Setup Flow (`setup.py` / `START-HERE.bat` / `setup.sh`)**:
   - Detects dual-GPU workstations with $\le 48\text{ GB}$ system RAM and automatically configures Dual-GPU Low-RAM mode (`--vram-experts`).
   - Sets `"gpu": [0, 1]` without `--layer-split`: GPU 0 computes and caches, while GPU 1 acts as the VRAM expert store (serving 21 whole layers via PCIe D2H to accelerate SSD paging).

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.