Pull requests / #1747

#1747 bench: community report, 2x PH402 SKU 200 (4x GP100, sm_60), Swift 1.5 IQ3_XXS at 1M, --pipeline-windows off vs on over four stages, engine 0.1.41 + patches

open · @ruibeikaa · 0 Kommentare · Auf GitHub

BenchmarksServer & APIMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows

Beschreibung

Results-only submission, no engine changes. Adds bench/results/2026-10-09-community-2x-ph402-sku200/ following docs/COMMUNITY_BENCHMARKS.md.

**Headline.** `--pipeline-windows 2` (#1656 + #1674) on a four-stage layer split of four GP100 dies raises decode by **+14.6% to +28.8%** on 9 prompts of 87 to 301,147 tokens. Some examples:
- 40.5 → 51.2 tok/s on a 256-token code answer.
- 47.6 → 61.3 on a 512-token one.
- 43.2 → 53.7 after a fresh 301,108-token read.
- 41.7 → 50.3 on a follow-up at 301K.

Prompt reads are unchanged by the flag. Replies are identical bit for bit in all three arms on all 9 prompts.

Hardware:
- **GPUs:** 2× NVIDIA PH402 SKU 200, i.e. four GP100 dies (sm_60, 48 SMs, 31.91 GiB each, TCC). Each die has a 140 W cap and an application clock of 1050 MHz.
- **Links:** the two dies of a board copy peer-to-peer at 73.0–73.2 GB/s; across boards 3.3 GB/s; host to device 3.3 GB/s per die.
- **Host:** Ryzen 9 9950X3D, 256 GiB DDR5, ASUS ProArt B850-CREATOR WIFI NEO, NVMe SSD.
- **OS:** Windows 11 (build 26300), driver 581.80, CUDA 12.9.

Software: source builds for sm_60 (`-DSTRATA_EXPERIMENTAL_SM60=ON`). The server is `serve/server.py` from v0.1.41.
- **next8:** v0.1.41 + #1674 + #1656 + the PH402 patch set. That set is #1424, #1441, #1660, #1368 and #1525, plus local sm_60 commits. The commit list is in the README.
- **next7:** the same patch set on v0.1.40.4, without pipelined windows.
- **next9:** next8 linked against ggml-cuda with a local sm_60 MMQ patch ([diff in my comment on #1639](https://github.com/Niko1221/Strata/issues/1639#issuecomment-6086238713)).

Model:
- [`ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF`](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF) IQ3_XXS. The dense projections and head were converted to Q8_0 with `tools/q8_dense_gguf.py` (#1424).
- Context 1,048,576 tokens (yarn ×4), int8 KV, layer split `12,24,36`, all 24,576 experts in VRAM.
- MTP `--spec 4 --spec-min-p 0.85 --mtp-window 8192`, greedy.

Configurations tested (every reply is hashed):
1. **Pipelined windows:** next7, next8 without the flag, and next8 with `--pipeline-windows 2`. One engine start each, the same 9 requests. These include a follow-up at 301K and a return to that conversation from the conversation cache.
2. **Prompt path:** next8 vs next9 (the MMQ patch), and `STRATA_PREFILL_PIPE` 768 / 290 / 450 on next9. Eight starts in bracketed order, 5 requests each (3 in the one start with the prefill timer).
3. **Draft floor:** `spec_min_p` 0.5 / 0.7 / 0.85 / 0.95 per request with pipelining on. 3 prompts × 2 rounds.

Other results:
- **MMQ patch:** the batched reads take 2.5–6.0% less time (bracketed next8 / next9 / next8 starts, one run each), and replies are identical.
- **`STRATA_PREFILL_PIPE` 768 → 290:**
  - The 1,158-token read goes from 4.627 / 4.632 s to 4.018 s.
  - The 34,835-token read goes from 48.967 / 49.195 s to 46.545 s.
  - 1,027 tokens appended to that conversation go from 4.339 / 4.348 s to 4.419 s (+1.7%).
  - The changed chunking changes the output bits.
- **Draft floor:** 0.85 is the fastest value on both code prompts.
- **Driver check:** `bench.py` (plan `ab`), run again on next9 at 290, reproduced that arm's 5 reply hashes.

Limitations:
- **Runs:** one run per arm for the pipelining A/B, 1–3 runs for the prompt-path arms, and 2 rounds for the floor sweep.
- **Baselines:** no stock v0.1.41 arm, and #1656 was not run without #1674.
- **Timers:** the prefill and decode timers were on in the three pipelining arms.
- **Not measured:** reply quality (replies were not graded, and every one ran to its cap), TTFT and peak memory.
- **Scope:** one model.
- **Hashes:** the sha256 of the converted Q8_0 shards and the pack/MTP preparation commands were not recorded. The README marks them.

Context:
- #1656 (@Cass67) and #1674 (@CC-David-CC): this adds four stages on CUDA sm_60 with reply hashes, at contexts up to 301K tokens. A trace breakdown is in [my comment on #1656](https://github.com/Niko1221/Strata/pull/1656#issuecomment-6086159617).
- #1441: `STRATA_PREFILL_PIPE`.
- #1639 (@thedrzinger, 5× P100): a PRMT dp4a emulation for the sm_60 MMQ path; [my reply there](https://github.com/Niko1221/Strata/issues/1639#issuecomment-6086238713) has the patch and its numbers.

Mehr auf der Site

Links zu Install, Modellen, Releases.