贡献 / #1747
#1747 bench: community report, 2x PH402 SKU 200 (4x GP100, sm_60), Swift 1.5 IQ3_XXS at 1M, --pipeline-windows off vs on over four stages, engine 0.1.41 + patches
open · @ruibeikaa · 0 评论 · 去 GitHub 看
BenchmarksServer & APIMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows
说明
Results-only submission, no engine changes. Adds bench/results/2026-10-09-community-2x-ph402-sku200/ following docs/COMMUNITY_BENCHMARKS.md. **Headline.** `--pipeline-windows 2` (#1656 + #1674) on a four-stage layer split of four GP100 dies raises decode by **+14.6% to +28.8%** on 9 prompts of 87 to 301,147 tokens. Some examples: - 40.5 → 51.2 tok/s on a 256-token code answer. - 47.6 → 61.3 on a 512-token one. - 43.2 → 53.7 after a fresh 301,108-token read. - 41.7 → 50.3 on a follow-up at 301K. Prompt reads are unchanged by the flag. Replies are identical bit for bit in all three arms on all 9 prompts. Hardware: - **GPUs:** 2× NVIDIA PH402 SKU 200, i.e. four GP100 dies (sm_60, 48 SMs, 31.91 GiB each, TCC). Each die has a 140 W cap and an application clock of 1050 MHz. - **Links:** the two dies of a board copy peer-to-peer at 73.0–73.2 GB/s; across boards 3.3 GB/s; host to device 3.3 GB/s per die. - **Host:** Ryzen 9 9950X3D, 256 GiB DDR5, ASUS ProArt B850-CREATOR WIFI NEO, NVMe SSD. - **OS:** Windows 11 (build 26300), driver 581.80, CUDA 12.9. Software: source builds for sm_60 (`-DSTRATA_EXPERIMENTAL_SM60=ON`). The server is `serve/server.py` from v0.1.41. - **next8:** v0.1.41 + #1674 + #1656 + the PH402 patch set. That set is #1424, #1441, #1660, #1368 and #1525, plus local sm_60 commits. The commit list is in the README. - **next7:** the same patch set on v0.1.40.4, without pipelined windows. - **next9:** next8 linked against ggml-cuda with a local sm_60 MMQ patch ([diff in my comment on #1639](https://github.com/Niko1221/Strata/issues/1639#issuecomment-6086238713)). Model: - [`ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF`](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF) IQ3_XXS. The dense projections and head were converted to Q8_0 with `tools/q8_dense_gguf.py` (#1424). - Context 1,048,576 tokens (yarn ×4), int8 KV, layer split `12,24,36`, all 24,576 experts in VRAM. - MTP `--spec 4 --spec-min-p 0.85 --mtp-window 8192`, greedy. Configurations tested (every reply is hashed): 1. **Pipelined windows:** next7, next8 without the flag, and next8 with `--pipeline-windows 2`. One engine start each, the same 9 requests. These include a follow-up at 301K and a return to that conversation from the conversation cache. 2. **Prompt path:** next8 vs next9 (the MMQ patch), and `STRATA_PREFILL_PIPE` 768 / 290 / 450 on next9. Eight starts in bracketed order, 5 requests each (3 in the one start with the prefill timer). 3. **Draft floor:** `spec_min_p` 0.5 / 0.7 / 0.85 / 0.95 per request with pipelining on. 3 prompts × 2 rounds. Other results: - **MMQ patch:** the batched reads take 2.5–6.0% less time (bracketed next8 / next9 / next8 starts, one run each), and replies are identical. - **`STRATA_PREFILL_PIPE` 768 → 290:** - The 1,158-token read goes from 4.627 / 4.632 s to 4.018 s. - The 34,835-token read goes from 48.967 / 49.195 s to 46.545 s. - 1,027 tokens appended to that conversation go from 4.339 / 4.348 s to 4.419 s (+1.7%). - The changed chunking changes the output bits. - **Draft floor:** 0.85 is the fastest value on both code prompts. - **Driver check:** `bench.py` (plan `ab`), run again on next9 at 290, reproduced that arm's 5 reply hashes. Limitations: - **Runs:** one run per arm for the pipelining A/B, 1–3 runs for the prompt-path arms, and 2 rounds for the floor sweep. - **Baselines:** no stock v0.1.41 arm, and #1656 was not run without #1674. - **Timers:** the prefill and decode timers were on in the three pipelining arms. - **Not measured:** reply quality (replies were not graded, and every one ran to its cap), TTFT and peak memory. - **Scope:** one model. - **Hashes:** the sha256 of the converted Q8_0 shards and the pack/MTP preparation commands were not recorded. The README marks them. Context: - #1656 (@Cass67) and #1674 (@CC-David-CC): this adds four stages on CUDA sm_60 with reply hashes, at contexts up to 301K tokens. A trace breakdown is in [my comment on #1656](https://github.com/Niko1221/Strata/pull/1656#issuecomment-6086159617). - #1441: `STRATA_PREFILL_PIPE`. - #1639 (@thedrzinger, 5× P100): a PRMT dp4a emulation for the sm_60 MMQ path; [my reply there](https://github.com/Niko1221/Strata/issues/1639#issuecomment-6086238713) has the patch and its numbers.
本站相关内容
相关页面的快捷入口。