Pull requests / #910

#910 Two-GPU decode: bounded attention merge (bitwise, opt-in) and pipelined PLE read-ahead / early chain (stacked on #905)

closed · @Hardin22 · 0 comments · View on GitHub

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsWindows

Description

Stacked on #905 (and through it on #859 and #876). Only the last two commits are new.

These are the last pieces of the two-GPU work. With them, main reaches what our fork does on the same PC.

1. **Decode attention: a merge bounded by the selection width** (`STRATA_ATTN_MERGE_V2=1`, opt-in since 9d27f7e: bitwise the same, but not yet measured apart from the other switches).
   - It is bitwise equal to the current merge. `attn_merge_parity` checks 768 cases (fp16 / int8 / q4_0 / k8v4 KV) on an RTX 5080 and an RTX 4060 Ti.
   - Faster or equal everywhere in the kernel bench, e.g. one query at full width on the 5080: 98 → 21 µs.
   - Off on HIP, where it has never run.

   Same commit: **`STRATA_MTP_KV=f16`** (opt-in), the draft layer's KV in f16. It changes only the drafts, never the verified tokens, and costs ~230 MiB of VRAM.
2. **Pipelined decode** (opt-in switches, only inside `--pipeline-windows`):
   - **`STRATA_PL_PLE_PREFETCH`** reads the next window's PLE rows ahead, as the chain's outputs land.
   - **`STRATA_PL_PLE_LATE`** gives each stage-0 window its own read handle, so `service` never blocks on another window's reads. Layer 0's experts are served while its flag waits for the rows.
   - **`STRATA_PL_EARLY_CHAIN`** launches the chain before the verdict's bookkeeping.

   The serial loop, `run()` and the batch windows are unchanged.

## Measured

Same setup as #905: 4060 Ti + 5080 split, resident RAM mode (#848), `--pipeline-windows 2 --adapt-async 1`, `--trim-stage-weights`, Swift 1.5 IQ3_XXS at 160K (q4_0 KV). Greedy, five prompts × two rounds, decode tok/s:

| | run 1 | run 2 |
|---|---|---|
| off (`STRATA_ATTN_MERGE_V2=0`, nothing else set) | 117.5 | 117.5 |
| on (`STRATA_PL_PLE_PREFETCH=1 STRATA_PL_PLE_LATE=1 STRATA_PL_EARLY_CHAIN=1 STRATA_MTP_KV=f16`, merge v2 by default) | **132.1** | **130.6** |

That is **+12%**; I have not split it per switch. For reference, our fork on the same prompts measured 131.7 and 133.7, so the stack is at parity with it.

## Checks

- **Exactness**: the three pipelined switches give the same greedy text on as off (4/4 prompts, 500 tokens, `STRATA_IQ_MT_MIN=1 --pcie-frac 0 --adapt-every 0`, blocking tier).
- **Stress**: with every switch on and `STRATA_PIPELINE_FORCE_MISS=2`, 11 varied requests gave no degenerate answer and no stall.

## The whole two-GPU stack, on this PC

RTX 4060 Ti 16 GB + RTX 5080 16 GB, i9-14900KF, 32 GB of RAM, Swift 1.5 IQ3_XXS at 160K:

| | decode tok/s |
|---|---|
| main today, the 5080 alone (what setup picks for 32 GB) | ~29 |
| main today, both cards (`--mmap-experts` once the file cache is warm) | ~64 |
| #848 (resident RAM mode on the split) | ~70 |
| + #859 pipelined windows | ~84 |
| + #876 / #905 async tier beside the pipeline, + #904, `--trim-stage-weights` | ~118 |
| **+ this** | **~131** |

The last two rows use a fine-tuned draft layer, `--spec-min-p 0.65` and Windows power throttling off; the rows above them use the stock draft layer. So read this table as an outline. Each PR has its own A/B with the same settings on both sides.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.