Pull requests / #521

#521 bench: add community RTX 4000 Ada 2-GPU vs 3-GPU layer-split results

closed · @lijieming · 0 コメント · GitHub で見る

BenchmarksMulti-GPUNVIDIA / CUDA

本文

Adds a community benchmark comparing two- and three-GPU Strata layer-split layouts on a Dell PowerEdge R7515 with RTX 4000 Ada cards.

Key observation: adding a 70 W RTX 4000 SFF Ada to two 130 W RTX 4000 Ada cards left median decode throughput essentially unchanged (~73 tok/s), but reduced the 16.6K fresh-prompt throughput from 2,110 tok/s to 779 tok/s on this workload.

The contribution includes:
- a concise hardware/software and model provenance summary;
- raw JSONL results for the 2-GPU and 3-GPU runs;
- aggregate timing/decode/prefill statistics;
- explicit notes on harness limitations and the two malformed/incorrectly-keyed smoke-test items.

This is a single-machine community result, not a general performance claim. Exact auto-selected per-GPU layer boundaries were not retained, and the detailed 3-GPU startup log was accidentally overwritten after the experiment, so no unsupported split-boundary or cache-count claims are made.

Source commit tested: `d9ab8435f654c368c586340d490915f6addf56a3`.

関連リンク

インストール・モデル・リリースへの站内リンク。