Issues / #1447
#1447 docs: MULTI_GPU.md says --pipeline-windows 2 and --adapt-async 1 exclude each other, but they combine
open · @adambenhassen · 0 comentários · No GitHub
BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows
Descrição
**Doc mismatch.** `docs/MULTI_GPU.md` (~line 242) says: > `--pipeline-windows` and `--adapt-async 1` exclude each other (the engine says so and keeps the pipeline). `src/program/generate.cpp` (~5970, v0.1.40.3) only turns async off beside `--pipeline-windows 2` when `STRATA_PIPELINE_ADAPT_ASYNC=0`. `DETAILS.md` and the `--pipeline-windows` section of MULTI_GPU.md also describe them working together. With both flags the startup log prints both `--pipeline-windows 2: two verifiers per stage...` and `asynchronous adaptive tier: ... rounds`, with no "is off" line. The line at ~242 looks stale. **Possibly related (question, not a bug report):** setup writes `--remote-expert-opt` into a two-GPU layer-split config. As far as I can read `generate.cpp` (~4725), it only does something with a helper cache (`--expert-cache-device1`), so on a plain split it is a no-op. But it is listed as a reason to turn off `--pipeline-windows` and `--adapt-async`. Is it intended for setup to add it on a split? **Setup** - 2x RTX 3090 24 GB (sm_86), 250 W limit each; GPU0 PCIe 4.0 x16, GPU1 PCIe 4.0 x4. No NVLink. - Ryzen 7 9800X3D, 32 GB RAM (30 GiB usable), Ubuntu 26.04.1, kernel 7.0.0, driver 610.43.02 - Strata v0.1.40.3 (`d5ea713`), source build with `setup.sh --build` in a CUDA 13.0.2 container, sm_86 - Model: Qwen3.8-Flash-Next GSQ-RCO IQ3_S (ISTA-DASLab), vision on GPU, 262,144 context, KV int8, `--spec 4 --mtp`, `--resident-experts` (15.95 GiB of experts page-locked), `--vram-reserve-mib 700` **Method.** For each arm: start the engine, send one small warm-up request, then wait until both GPUs are at or below 52 °C. Then: - six chat requests generating 1,500 tokens each (coding prompt, `reasoning_effort: low`, temperature 0.6, seeds 1-6); decode speed is `timings.predicted_per_second`, averaged over runs 2-6. - one cold ~19.9K-token prompt; prefill speed is `timings.prompt_per_second`. Arms with `--pipeline-windows 2` use `"layer_split": "29"` and drop `--remote-expert-opt`. Every other arm uses setup's config (`layer_split auto`, `--remote-expert-opt`). | Arm | Decode tok/s | Prefill tok/s | |---|---:|---:| | setup defaults (run twice) | 137.0 / 136.9 | 1798 | | `--adapt-async 1` (two runs) | 145.2 / 143.7 | 1800 / 1797 | | `--adapt-async 1`, no `--remote-expert-opt` | 145.0 | 1807 | | `STRATA_ADAPT_LAG=2` | 142.8 | 1800 | | `STRATA_EXCHANGE_ROTATE=1` | 139.6 | 1800 | | `STRATA_PF_FUSED=1` | 137.2 | 1913 | | `--pipeline-windows 2` | 136.7 | 2097 | | `STRATA_SPEC_COUPLED=1` | 137.6 | 1804 | | `STRATA_SPEC_PROB=1` | 137.2 | 1788 | | `STRATA_SPEC_COUPLED=1` + `STRATA_SPEC_GUMBEL=1` | 136.5 | 1805 | | async + lag2 + rotate + pf_fused | 145.0 | 1829 | | pw2 + lag2 + rotate + pf_fused | 135.9 | 2296 | | **pw2 + async + lag2 + rotate + pf_fused** | **142.0** | **2302** | - The last row gets most of async's decode gain plus pw2's +28% prefill. - No `timed out at layer` stalls in any arm. `STRATA_DF_BRANCH` was left off: on 0.1.39 plus the pipeline PRs it caused the layer-44 verify stall reported on #905. - On the same hardware, the setup defaults went from **120.4 tok/s** on v0.1.39 to **137.0 tok/s** on v0.1.40.3. - The sampled-drafting switches made no difference at temperature 0.6, which matches DETAILS.md.
No site
Links install, modelos, releases.