Pull requests / #1505
#1505 bench: community reports - 2x V100-SXM2 split (first measurement) + 2x RTX 2080 Ti 22GB
closed · @caolonghao · 0 评论 · 在 GitHub 查看
BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows
描述
Two results-only reports from one PC (2x Xeon Silver 4316, 125 GB RAM, NVIDIA driver 580.178.04, engine 0.1.40.2 source-built with the CUDA 12.8 toolkit, gcc 11.4, Ubuntu 22.04.5). No engine changes. Folders: ## bench/results/2026-10-08-community-2x-v100-sxm2/ **The first measured layer split over two V100** (docs/NVIDIA_V100.md: "A layer split over two V100 was not measured with this kernel") - 2x V100-SXM2-32GB on the experimental CUDA 12 engine (sm_70, STRATA_EXPERIMENTAL_SM60), Unsloth UD-IQ4_XS, 131,072 context, auto split K=25, the two caches holding 22,473 of 24,576 profiled pairs (~99.8%): - 4K / 32K fresh prompts: **1,277.8 / 2,355.3 tok/s prompt, 77.3 / 80.3 tok/s decode** (vs the single V100-SXM2 UD-IQ4_XS point in #902: decode 26.3 - the split is ~3x). - **Context ladder to 1M** (yarn up to 4.0): no VRAM wall - KV streaming keeps the expert caches at 99.7% coverage even at 1,048,576 (KV in 6.19 GiB of pinned RAM); decode 81.2 -> 70.6 across the ladder; the 512K->1M step at the same 450K prompt costs decode -5.6%, prompt speed unchanged. - **Concurrency**: queueing stays flat (~66 tok/s total); `"parallel": 2` adds no throughput on a split (batch windows carry no MTP drafts) but halves TTFT; `--batch 8 --batch-groups 2 --trim-stage-weights` reaches **117.7 tok/s total at 8 clients (+78%)**, 14.9 per request. - **Counter-example for a tip**: `--prefill auto:32768` (the +21-35% from #433/#440/#834, single-card data) measured **-32% at 32K** on this two-card split - one big chunk loses the cross-card chunk pipelining. - Operational note: a cold start on slow shared storage can exceed the server's 900 s READY watchdog; STRATA_ENGINE_READY_S is the workaround. ## bench/results/2026-10-08-community-2x-2080ti-22gb/ The same PC's 2x RTX 2080 Ti 22 GB pair (250 W limits - #1225's ran a 100 W cap), GSQ-RCO IQ3_XXS, same harness: **981.3 / 1,684.4 tok/s prompt, 84.1 / 75.1 tok/s decode** - vs #1225: 4K decode +59%, 32K prompt +253%. Both folders follow the COMMUNITY_BENCHMARKS.md template with per-run JSON, harness logs, the concurrency script, and the limitations spelled out (no needle/HumanEval, one machine, shared storage). Happy to adjust or split the PRs if you prefer one folder per PR. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。