Pull requests / #1636
#1636 batch-mtp: work with a layer split, +15% two-stream decode, one group under `--batch-groups auto`
open · @noon-at-cgn · 0 Kommentare · Auf GitHub
BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows
Beschreibung
## Title
Issue: Related: #1253 (ilumn, "Enable serial multi-GPU batch MTP": the same design, open, based on 82f46a8, includes #1242). #1413 (`--batch-groups G>1` + `--batch-mtp` is silently inert). #1242.
## Summary
**Why it matters:** on a layer split `--batch-mtp` is refused ("it is for one GPU (no layer split or helper) for now"), so two people using a two-GPU machine get one token per window each, while the flag gives +31% to +39% on one GPU (docs/BATCHING.md). This change lets `--batch-mtp` run on a layer split: each slot's drafter lives on the last stage's GPU (where the solo drafter and the head are), and a window carries two rows per slot (its token and its proposal) through every stage. On 0.1.41 it also needs one more fix: a layer split with `--batch` 2 or more now pipelines the slots in groups by default (`--batch-groups auto`), pipelined groups run no slot drafts (#1413), and the `--batch-mtp` gate ran before that default was resolved, so `--batch-mtp` stayed on while the groups became 2 and the drafts never ran. With `--batch-mtp` and no `--batch-groups`, `auto` now resolves to one group and the log says why. `--batch-groups N` (N above 1) or `--batch-groups auto` given on the command line is honoured and turns `--batch-mtp` off, as upstream does.
| Measured: 2 concurrent greedy decodes of 300 tokens, layer split `23`, `parallel 2`, `--spec 4`, 2x RTX 3080 20 GB (220 W cap), Xeon E5-2696 v4, UD-Q4_K_XL | aggregate tok/s (median of the two streams added) |
|---|---|
| **0.1.41** (our production stack on fb58e0d, binary strata-w10): `--batch-mtp`, one group (3 restarts, 60 measurements) | **89.2** (34.6 ms and 3.3 rows per window, 66-67% of proposals accepted) |
| same, no `--batch-mtp`, one serial group (2 restarts, 40) | 77.3 (23.7 ms, 2.0 rows) |
| same, no `--batch-mtp`, `--batch-groups` unset = the 0.1.41 default, two pipelined groups of one slot (1 restart, 20) | 77.7 |
| gain of `--batch-mtp` | +15.4% against serial (95% bootstrap interval +12.5% to +17.9%), +14.8% against the pipelined default; one stream alone is the same (81.7 and 82.0) |
| **0.1.40.3** + this change, UD-Q4_K_XL, without / with `--batch-mtp` (2 runs per arm) | 39.9-43.5 / 44.5-51.3 |
| **0.1.40.3**, Coder IQ1_M test model (all experts in VRAM), without / with | 84.0 / 101.2 (+20%), greedy text identical |
**Where upstream's default is better:** a new prompt of about 48k tokens arriving while one stream decodes was read at 2508 tok/s with the pipelined groups and 1217 tok/s in serial windows, and the decoding stream finished sooner (25.0 against 18.0 tok/s over its whole run). `--batch-mtp` does not change that. So `--batch-mtp` pays where concurrent decodes dominate; with long reads arriving beside decodes the 0.1.41 pipelined default can be the better choice. The first two 0.1.41 arms ran the asynchronous adaptive tier in batch windows, which is #1637 (the stack's next change; upstream main refuses it with batch slots); the pipelined path never adapts in windows (its tier runs in the solo path only), and that arm is our build started with the flags 0.1.41 resolves to, not an upstream binary.
## What changed
- `src/program/generate.cpp`: `--batch-mtp` is accepted with a layer split whose stages are on separate GPUs (still refused when two stages share a GPU, with `--batch-groups` above 1, and, new, with helper-GPU expert caches or `--remote-expert-opt` on a split, which have not been run). The slot drafters load from the last stage's slot sessions under that stage's device guard and bind to its weights and head; their residual buffers, the row copy (asynchronous, on the drafter's own stream) and the draft-K/V copies at admission and slot restore run on the drafter's device. Every stage's verifier gets the eight-row window size and the batch graph limit. The slot drafters' logits are not charged to CUDA0's cache sizing. The log says where the drafters are and what each costs (`--batch-mtp: 2 slot drafters on CUDA1 (50 MiB of private K/V state and buffers each ...)`); the `strata batch:` summary says how many proposals were accepted.
- `src/core/verify.cpp`, `include/strata/core/batch_rows.hpp` (new): the row-layout rules of `stage_batch` move into a pure function; grouped slot rows are allowed on a stage that hands rows on, and the hand-off bound is `kVerifyMaxT` rows (it was the slot count, which a slot's two rows exceed). CPU test `tests/core/batch_rows_test.cpp`.
- `include/strata/program/batch_groups.hpp` (new, pure) and `src/program/batch_groups_test.cpp` (CPU): how many groups `--batch` runs with (the 0.1.41 default, `auto`, an explicit number) and the one case where the default gives way (`--batch-mtp`). `generate.cpp` calls it before the `--batch-mtp` gate, so the gate sees the resolved number. Log line: `--batch-groups auto: 1 group of 2 slots - --batch-mtp does not run in pipelined groups; --batch-groups 2 would pipeline them and turn it off`.
- `docs/BATCHING.md`: the layer-split paragraph, the `--batch-groups auto` row, the limits list, the measurements above.
- Commits (all by noon-at-cgn): the layer-split work (2 commits, cherry-picks of the earlier branch), the gate fix, the helper-cache refusal with docs, the docs with the measurements.
## Extra Notes
**#1253 (ilumn / Alex, open, conflicts with main, based on 82f46a8, includes #1242): the same design, so please read both.** The same: slot drafters on the last stage's GPU bound to that stage's weights and head, serial windows (`--batch-groups 1`), eight-row windows through every stage with a bounded batch graph cache, `--batch-mtp` still refused when two stages share a GPU. What differs:
- #1253 has `OnDevice` guards in the destructors of `Verifier` and `MtpDrafter` and in `MtpDrafter::idle`, and takes the last stage's bind bytes (the drafters' logits and residual rows, `late_bind`) out of that stage's expert-cache room. This change has neither (it relies on `--vram-reserve-mib`); I did not copy them, since they are theirs.
- #1253 refuses `--pipeline-windows`, helper caches, `--remote-expert-opt` and peer devices in the gate. Here `--pipeline-windows` is already off with batch slots, `--peer-device` cannot be combined with a layer split, and helper caches on a split are refused.
- Here the row rules are a tested function (`batch_rows.hpp`); in #1253 they stay inline in `stage_batch`.
- #1253 carries a GPU lifecycle harness (`tools/batch_mtp_split_test.py`: staggered admission, solo promotion, seeded and greedy, mixed images) and measurements on 2x RTX 5090 with an IQ3_XXS model; this change has no GPU harness of its own, only the measurements above and a text-equality check (below).
- #1253 was written on 0.1.40.1 and its commands pass `--batch-groups 1`. Rebased on 0.1.41 its gate (`o.batch_groups > 1`, evaluated before `auto` is resolved) would see 1 and keep batch MTP on while the groups become 2. The gate fix here is what makes `--batch-mtp` work on 0.1.41 without that flag.
- #1253 includes #1242 (per-slot image positions, shared draft-head type); this change does not touch images, and the shared-head type line is already in main.
They cannot both be applied as they are: they conflict in `generate.cpp` (gate and drafter sites), `verify.cpp` (`stage_batch`) and `docs/BATCHING.md`. If #1253 goes first, what stays here is the gate fix and its test, `batch_rows.hpp` and its test, the accepted-proposals log line and the asynchronous row copy, and I will rebase onto it. If this goes first, #1253 still brings its harness, the destructor guards and the `late_bind` accounting. The gate fix needs neither. #1253's own table shows batch MTP slower than pipelined batching on 2x RTX 5090 (54.7 against 61.0 tok/s at 4000 experts per stage, 148.0 against 176.7 at 11000), and blange48's report on it found no gain on a 4-stage split (120/125/149/173/200 against 123/125/148/174/202 serial and 121/154/218/258/378 pipelined at 1/2/3/4/8 clients). Our two-stage result is for two stages and two slots. Related, not changed here: #1504 (long prompt reads on a layer split beside a decoding slot) concerns the read speed that the pipelined default wins in the table above.
**Text equality:** with the same expert cache on both arms (the slot drafters take VRAM, 4213 slots on CUDA1 in both after adjusting the reserve), `--pcie-frac 0 --adapt-every 0 --suffix-draft 0`, the greedy text with `--batch-mtp` equalled the text without it for 5 short prompts in all 6 comparisons (one restart per arm; the 25k-token prompt was not repeatable between two identical streams of the same arm). With different caches (4246 against 4213 slots) even the solo path differs on 5 of 6 prompts, so an unadjusted on/off comparison is confounded. This is text equality, not a bit-exactness proof, and says nothing about sampled output.
**Not tested:** HIP, SYCL, three or more stages (#1253's report found no gain at four), more than two slots, sampled decoding, images with `--batch-mtp` on a split (no per-slot image positions here: that is #1242), GPU parity of the split path beyond the text check above, `--batch-mtp` with helper caches (refused), any upstream-main binary (all 0.1.41 numbers are from our build). The new gate and log lines were read in the start log of the production engine on strata-w10 (this change plus the rest of our stack): the base config and `--batch-groups 2` (which printed `--batch-mtp is off: it does not combine with --batch-groups (pipelined slot groups) yet`); there was no CPU-only start of the engine.
Tests (CPU, `CUDA_VISIBLE_DEVICES=-1`): `batch_rows_test` and `batch_groups_test` (14 checks) pass; `ctest` on the whole tree: 52 pass, 56 fail with no GPU (the same set on this branch and #1637). On GPU 0 with about 1 GiB free: `verify_parity`, `verify_batch_parity`, `rope_parity`, `qsa_parity`, `kv_stream_parity`, `kv_hybrid_parity` pass; `gdn_rec_parity` stops with `cudaMalloc: out of memory` (another process holds that GPU's memory). `python -m unittest discover -s serve`: 593 tests OK (11 skipped).
Context: this change is part of the setup measured in the community benchmark #1640.
Mehr auf der Site
Links zu Install, Modellen, Releases.