Pull requests / #151

#151 Add secondary GPU expert store (--vram-experts) and low-RAM multi-GPU setup

closed · @lukmanfauzie · 0 comentários · No GitHub

Setup & installMulti-GPUAMD / HIPModels & quantsWindows

Descrição

Replaces and supersedes #111. Rebased into a single clean commit on top of latest `main` (`v0.1.24`), addressing all review feedback:

1. **Native router dynamic expert count + batched multi-token launch**:
   - Upstream v0.1.24's batched `native_router_top10_multi` now takes `n_expert` row stride dynamically.
   - Seamlessly handles both standard 512-expert models and 256-expert pruned models (Coder).

2. **MTP drafter**:
   - Preserves upstream v0.1.24's bulk `FILE*` loader in `mtp.cpp` while taking probed draft expert count.

3. **Native pack `--mmap-experts` support**:
   - Native IQ packs shipping `experts.bin` (via `tools/iq_pack.py --experts-bin`) map directly. `FileExpertSource::offset_of` indexes blobs using `ExpertLayout::blob_offset`, addressing both uniform and per-layer variable layouts.

4. **Server tests & telemetry preserved**:
   - Restored all 8 server tests; the entire test suite (38/38 server tests, 73/73 unit/integration tests) passes cleanly.
   - Live prompt throughput indicators (`prompt_tok_s`, `prompt_eta_s`) feed to `app.js` via upstream's `PP` pipeline.

5. **Dual-GPU Low-RAM Setup Flow (`setup.py` / `START-HERE.bat` / `setup.sh`)**:
   - Detects dual-GPU workstations with $\le 48\text{ GB}$ system RAM and automatically configures Dual-GPU Low-RAM mode (`--vram-experts`).
   - Sets `"gpu": [0, 1]` without `--layer-split`: GPU 0 computes and caches, while GPU 1 acts as the VRAM expert store (serving 21 whole layers via PCIe D2H to accelerate SSD paging).

No site

Links install, modelos, releases.