Pull requests / #1251

#1251 --peer-device prompt share without P2P: host route, per-token sums, FP16 transfers (2.1-2.5x prompts on a no-P2P pair)

open · @guthirry · 0 comments · View on GitHub

BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsSecurity

Description

## What

`--peer-device`'s prompt share for GPU pairs **without P2P** (most GeForce pairs). Before this PR, `set_peer` refused such a peer and the prompt path stayed on the primary GPU. Now:

1. **The host route.** Without P2P, the share uses the route the layer split's helper already uses: mapped pinned host buffers, read by copy kernels, with the compact group buffers. `STRATA_PF_PEER_HOST=0` restores the refusal.
2. **Sums instead of rows.** The peer adds its experts' rows into one weighted sum per token and sends back T x N floats instead of its rows. Its rows are up to K = 10 times more data. `moe_combine_peer` adds the sums to the primary's own rows. `STRATA_PF_PEER_SUMS=0` sends rows.
3. **FP16 both ways.** The MoE input goes out as the existing `mixed_h` image, and the sums come back as FP16, saturated at ±65504. `STRATA_PF_PEER_F16=0` uses FP32.
4. **One launch per group** for the sums (`peer_gather_add`), bitwise the same as one launch per expert. `STRATA_PF_PEER_GATHER=0` launches per expert.
5. **The hand-off copies on a stream of their own**, so the primary's own experts start without waiting for the copies over its link. `STRATA_PF_PEER_XSTREAM=0` puts them on the compute stream.

The P2P route is unchanged: everything new sits behind `!p2p`.

**Credit.** The design of 2-4 is eddoursul/Strata's second-GPU prompt path (custom branch commits f8de703, 2acfac4 and e8ade66; `moe_scatter_add` / `moe_gather_add` / `sums_to_f16`). The commits that carry it list Eddoursul as co-author.

## Measured

**Setup:** RTX 3090 (PCIe 4.0 x4, behind the chipset; the PCIe probe reads 6.7 GB/s) and RTX 3080 Ti (x16). No P2P. i5-14600K, 93 GiB RAM, UD-Q4_K_XL. Both arms ran `--peer-device 1 --peer-reserve-mib 3000 --prefill auto:32768 --kv int8 --kv-resident 32768 --spec 4 --pcie-frac 0 --peer-adapt-swaps 384`, with `STRATA_PF_PEER_STREAM=0.6 STRATA_PF_PEER_RING=160`.

**Method:** the same OpenAI requests through `serve/server.py`, greedy, one at a time. Both arms use this branch's build, with the share off (`STRATA_PF_PEER_HOST=0`) or on.

Prompt tok/s:

| Prompt | Share off (= main) | This PR | Change |
|---|---:|---:|---:|
| 3.4K tokens | 332 | 815 | 2.5x |
| 8K tokens | 763 | 1,636 | 2.1x |
| 100K tokens (4 chunks) | 1,754 | 1,926 | +10% |

Decode tok/s:

| Workload | Share off | This PR |
|---|---:|---:|
| 8K repeat, after warm-up | 120-123 | 120-123 |
| 3,221-token edit | 98 | 97 |
| 100K | 58 | 58 |

**Where the time went.** `STRATA_PREFILL_TIMING=1` with the share off shows 70% of an 8K prompt's GPU time in `wait copy`: the primary streaming experts over its x4 link. After this PR, the remaining wait is the peer's hand-off. The 3080 Ti is idle about 60% of the time, because the hand-off is serial per layer and the dense steps run over the whole chunk. Overlapping it the way the fork does would need the dense steps in sub-chunks; that is not part of this PR.

**Settings these numbers used**, beyond the defaults (all opt-in already):
- `STRATA_PF_PEER_STREAM=0.6`. The default 0.35 was measured on P2P pairs. Here at 8K: 0.35 gives 1,140, 0.6 gives 1,284, 1.0 gives 1,189 (before 3-5).
- `STRATA_PF_PEER_RING=160` with `--peer-reserve-mib 3000`. The 48-slot ring holds a third of a layer's peer-streamed experts at 0.6. At 8K: 48 slots give 1,485, 160 give 1,671.
- UD-Q4_K_XL needs the build with `-DSTRATA_MMQ_KQUANTS=ON`, because the peer share needs MMQ on every layer. Without it, `set_peer` says "needs the MMQ prompt path" and the prompt stays on the primary, as before.

**For reference:** eddoursul/Strata's custom branch on the same box, same requests: 974 / 2,082 / 2,228 prompt tok/s.

## Exactness

- **Byte-identical where tested.** The first two requests of every process (8K: 256 tokens; 3.4K: the 3,221-token edit) produced the same answers, byte for byte, in all arms: share off twice, share on, and share on with FP32 transfers. The edit answer is also the exact expected file.
- **The 12-excerpt check is inconclusive.** We ran 12 different 4-6K-token code excerpts with 128 greedy tokens each, after those first requests. These answers are not reproducible between two processes of the unchanged build either: share off vs share off gave 1 of 12 identical, with the first difference as early as the 2nd character. Share off vs share on gave 1 of 12, and FP16 vs FP32 gave 1-2 of 12. So this check cannot separate the PR's rounding (FP16 transfers, the sums' order) from main's own run-to-run variation. All answers read as correct summaries of their excerpts.
- **No teacher-forced comparison.** `generate` does not run the peer prompt share (`set_peer` is on the serve path), so we could not do a teacher-forced logit comparison.

## Not in this PR

- A share chosen by chunk size made 3.4K prompts faster than the fork (1,029 tok/s), but it lowered the following edit's decode by 16% through `DraftPolicy`. That is reported separately (https://github.com/Niko1221/Strata/issues/1252) and the rule is left out.
- A decode prefetch for the peer tier (the fork's f4ee9b5 idea) was tried and broke even on this box, so it is not included.

## Testing

- Built with CUDA 13.1, sm_86, `-DSTRATA_MMQ_KQUANTS=ON`. No warnings in the touched files.
- No new unit tests. The new kernels run in the serve measurements above, and the env switches give A/B paths for each step.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.