Pull requests / #389
#389 bench: community report, 2x TITAN RTX (sm_75), IQ3_S at 262K
closed · @shrisha108 · 0 commentaires · Sur GitHub
BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsDocumentationLinux
Description
Community benchmark per `docs/COMMUNITY_BENCHMARKS.md`. **Hardware:** 2x NVIDIA TITAN RTX 24 GB (Turing, sm_75), Xeon E5-2696 v4 (no AVX-512), 125 GiB RAM, NVMe 1 TB. PCIe x16 gen 3 under load. `nvidia-smi topo -m` reports `NV2`, but the engine does not use it, so the same numbers should be expected without a bridge. **Software:** AlmaLinux 9.8, kernel 5.14, driver 615.71.09, CUDA 12.9, Strata at commit `9259cad` (engine 0.1.31) built from source with `CMAKE_CUDA_ARCHITECTURES=75`. **Model:** original Flash-Next IQ3_S (3.5 bpw), two GGUF shards plus the MTP draft layer packed by setup. **Configuration:** `--expert-cache auto --prefill auto --spec 4 --mtp ... --max-context 262144 --kv int8 --kv-resident 65536 --ple-io ram --layer-split auto`. **Results** (three runs per length, prompt tokens; `benchmark.py` reproduces all nine): | Prompt | Prompt t/s | Decode t/s | TTFT | | --- | --- | --- | --- | | 4,096 | 840.9 / 865.4 / 864.7 | 67.3 / 79.5 / 68.1 | 3.54-3.64 s | | 32,768 | 1619.9 / 1611.8 / 1603.2 | 59.5 / 62.6 / 57.9 | 15.32-15.48 s | | 128,000 | 1786.2 / 1782.9 / 1784.6 | 61.6 / 59.1 / 62.6 | 54.28-54.40 s | Expert cache hit rate 99.0-99.7%. `needle_bench.py --lengths 32k,128k --depths 10,50,90`: 6 of 6 found, no misses. Two notes that may save someone else a purchase or a config change: - **NVLink is not used.** The engine deliberately avoids peer-to-peer and moves activations through pinned RAM once per verify window, not twice per layer (`docs/MULTI_GPU.md`). Do not expect a gain from bridging two consumer cards. - **`--ple-io ram` measured the same as the default `direct`** (1596.0 vs 1595.9 t/s on the same 176,460-token prompt) while holding 27.1 GiB more resident. Worth it only on a machine with RAM to spare; on a smaller one the default is better. Not measured: vision (the ready-made encoder has no sm_75 code) and the low-RAM variant. Credentials are stripped from the published config, and `benchmark.py` reads the key from `STRATA_API_KEY`. One more thing worth recording from the submitter, since the numbers alone do not say why this was worth the weekend: > I did not expect any of this to be possible. The whole point of the exercise was to see whether a 125B MoE could run at all on two consumer Turing cards, and it does - 60 tok/s of decode with the model's full 262,144-token window, on hardware that was never sold for this. Thank you.
Sur le site
Liens install, modèles, releases.