Pull requests / #1505

#1505 bench: community reports - 2x V100-SXM2 split (first measurement) + 2x RTX 2080 Ti 22GB

closed · @caolonghao · 0 comentarios · En GitHub

BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows

Descripción

Two results-only reports from one PC (2x Xeon Silver 4316, 125 GB RAM, NVIDIA driver 580.178.04,
engine 0.1.40.2 source-built with the CUDA 12.8 toolkit, gcc 11.4, Ubuntu 22.04.5). No engine
changes. Folders:

## bench/results/2026-10-08-community-2x-v100-sxm2/

**The first measured layer split over two V100** (docs/NVIDIA_V100.md: "A layer split over two V100
was not measured with this kernel") - 2x V100-SXM2-32GB on the experimental CUDA 12 engine
(sm_70, STRATA_EXPERIMENTAL_SM60), Unsloth UD-IQ4_XS, 131,072 context, auto split K=25, the two
caches holding 22,473 of 24,576 profiled pairs (~99.8%):

- 4K / 32K fresh prompts: **1,277.8 / 2,355.3 tok/s prompt, 77.3 / 80.3 tok/s decode** (vs the
  single V100-SXM2 UD-IQ4_XS point in #902: decode 26.3 - the split is ~3x).
- **Context ladder to 1M** (yarn up to 4.0): no VRAM wall - KV streaming keeps the expert caches
  at 99.7% coverage even at 1,048,576 (KV in 6.19 GiB of pinned RAM); decode 81.2 -> 70.6 across
  the ladder; the 512K->1M step at the same 450K prompt costs decode -5.6%, prompt speed unchanged.
- **Concurrency**: queueing stays flat (~66 tok/s total); `"parallel": 2` adds no throughput on a
  split (batch windows carry no MTP drafts) but halves TTFT; `--batch 8 --batch-groups 2
  --trim-stage-weights` reaches **117.7 tok/s total at 8 clients (+78%)**, 14.9 per request.
- **Counter-example for a tip**: `--prefill auto:32768` (the +21-35% from #433/#440/#834,
  single-card data) measured **-32% at 32K** on this two-card split - one big chunk loses the
  cross-card chunk pipelining.
- Operational note: a cold start on slow shared storage can exceed the server's 900 s READY
  watchdog; STRATA_ENGINE_READY_S is the workaround.

## bench/results/2026-10-08-community-2x-2080ti-22gb/

The same PC's 2x RTX 2080 Ti 22 GB pair (250 W limits - #1225's ran a 100 W cap), GSQ-RCO IQ3_XXS,
same harness: **981.3 / 1,684.4 tok/s prompt, 84.1 / 75.1 tok/s decode** - vs #1225: 4K decode
+59%, 32K prompt +253%.

Both folders follow the COMMUNITY_BENCHMARKS.md template with per-run JSON, harness logs, the
concurrency script, and the limitations spelled out (no needle/HumanEval, one machine, shared
storage). Happy to adjust or split the PRs if you prefer one folder per PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

En el sitio

Enlaces a install, modelos, releases.