Pull requests / #151
#151 Add secondary GPU expert store (--vram-experts) and low-RAM multi-GPU setup
closed · @lukmanfauzie · 0 commentaires · Sur GitHub
Setup & installMulti-GPUAMD / HIPModels & quantsWindows
Description
Replaces and supersedes #111. Rebased into a single clean commit on top of latest `main` (`v0.1.24`), addressing all review feedback:
1. **Native router dynamic expert count + batched multi-token launch**:
- Upstream v0.1.24's batched `native_router_top10_multi` now takes `n_expert` row stride dynamically.
- Seamlessly handles both standard 512-expert models and 256-expert pruned models (Coder).
2. **MTP drafter**:
- Preserves upstream v0.1.24's bulk `FILE*` loader in `mtp.cpp` while taking probed draft expert count.
3. **Native pack `--mmap-experts` support**:
- Native IQ packs shipping `experts.bin` (via `tools/iq_pack.py --experts-bin`) map directly. `FileExpertSource::offset_of` indexes blobs using `ExpertLayout::blob_offset`, addressing both uniform and per-layer variable layouts.
4. **Server tests & telemetry preserved**:
- Restored all 8 server tests; the entire test suite (38/38 server tests, 73/73 unit/integration tests) passes cleanly.
- Live prompt throughput indicators (`prompt_tok_s`, `prompt_eta_s`) feed to `app.js` via upstream's `PP` pipeline.
5. **Dual-GPU Low-RAM Setup Flow (`setup.py` / `START-HERE.bat` / `setup.sh`)**:
- Detects dual-GPU workstations with $\le 48\text{ GB}$ system RAM and automatically configures Dual-GPU Low-RAM mode (`--vram-experts`).
- Sets `"gpu": [0, 1]` without `--layer-split`: GPU 0 computes and caches, while GPU 1 acts as the VRAM expert store (serving 21 whole layers via PCIe D2H to accelerate SSD paging).Sur le site
Liens install, modèles, releases.