Pull requests / #1193

#1193 bench: community report, 2x TITAN RTX (sm_75), IQ3_S at 262K

closed · @shrisha108 · 0 コメント · GitHub で見る

BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentation

本文

Re-cut of #389 after the main force-push closed the original branch. Same measurements, same files, rebased onto current main alongside the MI50 and Arc Pro B60 reports.

Community benchmark per docs/COMMUNITY_BENCHMARKS.md: HP Z440, 2x TITAN RTX 24 GB (sm_75), Xeon E5-2696 v4 (no AVX-512), 128 GiB RAM, PCIe x16 gen 3, NVLink present but unused. Strata 0.1.31 built from source for sm_75. IQ3_S at the model native 262,144-token window.

| Prompt | Prompt t/s (3 runs) | Decode t/s (3 runs) | TTFT |
|---|---|---|---|
| 4,096 | 840.9 / 865.4 / 864.7 | 67.3 / 79.5 / 68.1 | 3.54-3.64 s |
| 32,768 | 1619.9 / 1611.8 / 1603.2 | 59.5 / 62.6 / 57.9 | 15.32-15.48 s |
| 128,000 | 1786.2 / 1782.9 / 1784.6 | 61.6 / 59.1 / 62.6 | 54.28-54.40 s |

Expert cache 99.0-99.7% hit. needle_bench 6/6 at 32k and 128k.

Notes: NVLink is not used by the engine; --ple-io ram measured equal to the default while holding 27 GiB more resident. Credentials stripped from the published config.

関連リンク

インストール・モデル・リリースへの站内リンク。