Pull requests / #1508

#1508 bench: community report, 2x Tesla V100-PCIE-32GB layer split (first two-V100 measurement), UD-IQ4_XS, context ladder to 1M, concurrency

open · @caolonghao · 0 comments · View on GitHub

BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows

Description

Results-only PR: one report folder, `bench/results/2026-10-08-community-2x-v100-pcie-32gb/` (a second PR will carry the same PC's 2x RTX 2080 Ti 22 GB pair).

**The first measured layer split over two V100** - docs/NVIDIA_V100.md: "A layer split over two V100 was not measured with this kernel" (a V100 + RTX 4070 split exists in #850; this is the first with two V100). Notably on **PCIe cards with no NVLink** (PCI device ID `10DE:1DB5`, all NVLink links inactive; the driver's "SXM2" name string is a VBIOS label also seen in NVIDIA forum reports of PCIe V100s) - the split rides PCIe Gen3 x16 (13.1 GB/s probe) and pinned host RAM only.

Hardware: 2x Xeon Silver 4316, 125 GB RAM, Ubuntu 22.04.5, driver 580.178.04, engine 0.1.40.2 source-built with the CUDA 12.8 toolkit (sm_70, `STRATA_EXPERIMENTAL_SM60`), Unsloth UD-IQ4_XS, 131,072 context, auto split K=25 (caches 22,473/24,576 pairs, ~99.8%). Same harness as #1225/#995 (`benchmark.py` unmodified, fresh nonce prompts, `reused: 0`, 3 runs per cell).

- 4K / 32K fresh prompts: **1,277.8 / 2,355.3 tok/s prompt, 77.3 / 80.3 tok/s decode**. Against the single V100-SXM2 UD-IQ4_XS point in #902 (decode 26.3; a different host/OS/engine/KV, one run there): ~3x decode with ~2.2x the experts resident.
- **Context ladder to 1M** (yarn up to 4.0): no VRAM wall - KV streaming keeps the expert caches at 99.7% coverage even at 1,048,576 (KV in 6.19 GiB of pinned RAM); decode 81.2 -> 70.6 across the ladder (the 131K/128K rung used `--prefill auto:32768`; decode was unchanged by that flag); the 512K->1M step at the same 450K prompt costs decode -5.6%, prompt speed unchanged.
- **Concurrency**: queueing flat (~66 tok/s total); `"parallel": 2` adds no throughput on a split (batch windows carry no MTP drafts) but improves TTFT (2.4 s -> 1.3 s at C=2); `--batch 8 --batch-groups 2 --trim-stage-weights` reaches **117.7 tok/s total at 8 clients (+78%)**, 14.9 per request.
- **Counter-example for a tip**: `--prefill auto:32768` (the +21-35% from #433/#440, single-card data) measured **-32% at 32K** on this two-card split; #834 saw the same reversal on a 32 GB 4090.
- Operational: a cold start on slow shared storage exceeded the server's 900 s READY watchdog; `STRATA_ENGINE_READY_S` is the workaround.

Limitations are spelled out in the README (no needle/HumanEval, one machine, shared storage, rope quality past 262K untested). Raw dumps and logs are pre-trimmed per the repository's practice (aggregates carry every number); the COMMUNITY.md index rows are left to the maintainers per COMMUNITY_BENCHMARKS.md.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.