Pull requests / #1154

#1154 decode: support a three-GPU two-window pipeline on 3 P4s

open · @CC-David-CC · 0 commentaires · Sur GitHub

BenchmarksNVIDIA / CUDAModels & quantsDocumentationWindows

Description

Three split GPUs currently disable `--pipeline-windows 2`. This change lets one conversation use two speculative windows across three stages, including an opt-in early start on the first stage. Every prefix stage keeps its own rollback snapshot, and commits stop at EOS or the output budget. Pipelined verifiers leave single-token commits to the host verdict.

The R730 recipe uses three Tesla P4s with x16/x8/x16 links: `gpu: [0,1,2]`, `layer_split: "20,28"` (20/8/20 layers), `--pipeline-windows 2`, and `STRATA_PIPELINE_PREFIX_EARLY=1`. All cards read one host expert backing. Optional per-device prefill helper shares and helper selection let the x16 cards help the narrow middle stage. The CUDA driver resolver also retains compatibility with CUDA 12.0 used on this Pascal machine.

Validation and measurements are recorded in `bench/results/2026-10-06-r730-three-p4/README.md`. The preceding patched v0.1.40 build reached median 26.88 tok/s on code and 18.12 tok/s on prose with Q8_0 weights, FP16 KV, 4,096 input tokens and a 512-token output cap. These are hardware results, not a stock-versus-patch speedup claim. The comparison to patched v0.1.39 was +2.9%/+3.4% decode, with prefill essentially unchanged at 147-149 tok/s.

This PR excludes the experimental shared-arena reuse and late PLE-residency utilities. Q4/IQ3_S, other hardware and sustained concurrent serving have not been validated on this final branch. The three-stage path remains opt-in.

Exact clean branch checks: Pascal Release build and commit_limit_test passed; all 30 requests matched serial tokens and all-stage recurrent/attention state; all six forced-rollback requests exercised rollback; five measured code outputs passed 1,005 functional cases each.

Sur le site

Liens install, modèles, releases.