Pull requests / #910
#910 Two-GPU decode: bounded attention merge (bitwise, opt-in) and pipelined PLE read-ahead / early chain (stacked on #905)
closed · @Hardin22 · 0 commentaires · Sur GitHub
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsWindows
Description
Stacked on #905 (and through it on #859 and #876). Only the last two commits are new. These are the last pieces of the two-GPU work. With them, main reaches what our fork does on the same PC. 1. **Decode attention: a merge bounded by the selection width** (`STRATA_ATTN_MERGE_V2=1`, opt-in since 9d27f7e: bitwise the same, but not yet measured apart from the other switches). - It is bitwise equal to the current merge. `attn_merge_parity` checks 768 cases (fp16 / int8 / q4_0 / k8v4 KV) on an RTX 5080 and an RTX 4060 Ti. - Faster or equal everywhere in the kernel bench, e.g. one query at full width on the 5080: 98 → 21 µs. - Off on HIP, where it has never run. Same commit: **`STRATA_MTP_KV=f16`** (opt-in), the draft layer's KV in f16. It changes only the drafts, never the verified tokens, and costs ~230 MiB of VRAM. 2. **Pipelined decode** (opt-in switches, only inside `--pipeline-windows`): - **`STRATA_PL_PLE_PREFETCH`** reads the next window's PLE rows ahead, as the chain's outputs land. - **`STRATA_PL_PLE_LATE`** gives each stage-0 window its own read handle, so `service` never blocks on another window's reads. Layer 0's experts are served while its flag waits for the rows. - **`STRATA_PL_EARLY_CHAIN`** launches the chain before the verdict's bookkeeping. The serial loop, `run()` and the batch windows are unchanged. ## Measured Same setup as #905: 4060 Ti + 5080 split, resident RAM mode (#848), `--pipeline-windows 2 --adapt-async 1`, `--trim-stage-weights`, Swift 1.5 IQ3_XXS at 160K (q4_0 KV). Greedy, five prompts × two rounds, decode tok/s: | | run 1 | run 2 | |---|---|---| | off (`STRATA_ATTN_MERGE_V2=0`, nothing else set) | 117.5 | 117.5 | | on (`STRATA_PL_PLE_PREFETCH=1 STRATA_PL_PLE_LATE=1 STRATA_PL_EARLY_CHAIN=1 STRATA_MTP_KV=f16`, merge v2 by default) | **132.1** | **130.6** | That is **+12%**; I have not split it per switch. For reference, our fork on the same prompts measured 131.7 and 133.7, so the stack is at parity with it. ## Checks - **Exactness**: the three pipelined switches give the same greedy text on as off (4/4 prompts, 500 tokens, `STRATA_IQ_MT_MIN=1 --pcie-frac 0 --adapt-every 0`, blocking tier). - **Stress**: with every switch on and `STRATA_PIPELINE_FORCE_MISS=2`, 11 varied requests gave no degenerate answer and no stall. ## The whole two-GPU stack, on this PC RTX 4060 Ti 16 GB + RTX 5080 16 GB, i9-14900KF, 32 GB of RAM, Swift 1.5 IQ3_XXS at 160K: | | decode tok/s | |---|---| | main today, the 5080 alone (what setup picks for 32 GB) | ~29 | | main today, both cards (`--mmap-experts` once the file cache is warm) | ~64 | | #848 (resident RAM mode on the split) | ~70 | | + #859 pipelined windows | ~84 | | + #876 / #905 async tier beside the pipeline, + #904, `--trim-stage-weights` | ~118 | | **+ this** | **~131** | The last two rows use a fine-tuned draft layer, `--spec-min-p 0.65` and Windows power throttling off; the rows above them use the stock draft layer. So read this table as an outline. Each PR has its own A/B with the same settings on both sides.
Sur le site
Liens install, modèles, releases.