Pull requests / #492

#492 Heterogeneous multi-GPU roles: whole model on one GPU, the MTP draft head on the other

closed · @hireymage · 0 commentaires · Sur GitHub

Multi-GPUNVIDIA / CUDAModels & quantsDocumentation

Description

Closes/addresses discussion in #490.

## What this adds

**Heterogeneous multi-GPU roles** — a different way to use two (not necessarily equal) GPUs than `--split-device`: the whole model (KV + verify) stays on the **main** GPU, only the **MTP draft head** runs on the second one. No P2P required, no layer split. The drafter is tiny next to the main model, so moving it off the main GPU returns its VRAM (~870 MiB + the draft head) to the main GPU's expert cache, while the drafter gets its own device entirely.

New (all opt-in, **no behavior change when none are passed**):

- `--main-device N --draft-device M` — explicit role placement; both directions work. The drafter's K/V, ring restore, gr workspace and graph captures are re-homed on the draft device; the verify window's hand-off rides a mapped host mirror, so no peer access is ever assumed;
- `--auto-roles` — pick the roles from the capability records (`main` = device 0, draft = the best free remaining; refuses when < 2 CUDA devices or no `--mtp` is loaded);
- `--draft-prefill-parallel` — the prompt fill queues and overlaps the decode instead of blocking it;
- `--draft-chain-batch` — the draft chain's step graphs launch back to back with ONE wait per round instead of one per step (drafting 4.63 → 4.14 ms/round on the source rig, ~11 %);
- `per-device DeviceCaps` discovery + the peer matrix in `strata-device`, and a per-role VRAM report after binding, so the trade is visible at start.

## The discipline: placement never changes the output, only the measured time

All claims are on greedy token streams compared bit-exactly across repeated runs, compared **under equal expert residency** (equal auto slot counts, or a forced `--expert-cache N` shared by both runs) — because when the roles split, the drafter's VRAM frees up on the main GPU, the expert cache there grows, and a different expert-cache size legitimately decodes differently today. Under that rule the roles 0→1, 1→0, reversed, batched and pipelined all reproduce the same-device stream token for token.

The plan, the measurements and the step record live in `docs/HETERO_MGPU_PLAN.md` (sections 10–12 hold the Fase-7/8/9/12/13 numbers and the gate rule).

## How to reproduce on your own rig

```
tools/hetero-test.sh <build-dir> "<your model-load flags>" /tmp/het 60 2
```
runs the variant matrix (single- and two-GPU rigs; the harness detects the rig itself), extracts the greedy stream md5s, picks bit-exactness references by residency, and writes `/tmp/het/report.md` — the full run table, verdicts and the residency-dividend note. Any report pasted into #490 is exactly the evidence this feature still needs from non-Pascal rigs.

Source-rig validation (2× GTX 1080 Ti, cc 6.1, 125B Q2_0 pack, 16 runs): **20 PASS / 0 FAIL**, stream md5 clusters by residency (V0=V1f=V2 `c1c2827c7571`; V1=V3–V6 `dc295284ab30`), V1f vs V0 bit-identical — the placement-exactness proof.

## Review notes

- This PR deliberately carries **no sm_61/Pascal build work** — it was developed on that rig (and the full branch is on my fork as `hetero-roles`), but only the role/flags/harness features are proposed here.
- Default-off everywhere; the existing `--split-device` path is untouched.

Sur le site

Liens install, modèles, releases.