Pull requests / #1429
#1429 bench: 2x TITAN RTX (sm_75) at Strata v0.1.40.3, IQ3_S at 262K
open · @shrisha108 · 0 comments · View on GitHub
BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentationLinux
Description
Community benchmark per `docs/COMMUNITY_BENCHMARKS.md`, re-measured on the current release tag. **Hardware:** 2x NVIDIA TITAN RTX 24 GB (Turing, sm_75), `nvidia-smi topo -m` reports `NV2` but the engine does not use it, so the same numbers should be expected without a bridge. Xeon E5-2696 v4 (no AVX-512), 125 GiB RAM, PCIe x16 gen 3 under load. **Software:** AlmaLinux 9.8, driver 615.71.09, CUDA 12.9, Strata tag **`v0.1.40.3`** (commit `d5ea713`) built from source with `CMAKE_CUDA_ARCHITECTURES=75`. **Configuration:** `--expert-cache auto --prefill auto --spec 4 --mtp ... --max-context 262144 --kv int8 --kv-resident 65536 --vision --vram-reserve-mib 700 --ple-io ram --layer-split auto --reasoning-budget-tokens 12000`. **Results** (three runs per length, prompt tokens; `benchmark.py` reproduces all nine): | Prompt | Prompt t/s | Decode t/s | TTFT | | --- | --- | --- | --- | | 4,096 | 831.8 / 876.3 / 868.9 | 65.9 / 70.9 / 71.5 | 3.50-3.69 s | | 32,768 | 1552.1 / 1548.9 / 1540.3 | 77.3 / 79.2 / 77.0 | 15.97-16.10 s | | 128,000 | 1636.7 / 1626.9 / 1621.7 | 71.1 / 73.8 / 70.7 | 59.13-59.69 s | 8,649 experts cached (16.00 GiB of VRAM), expert cache 95-99% hit. `needle_bench.py --lengths 32k,128k --depths 10,50,90`: **6 of 6 found**, no misses, at 125,918-token actual prompts. Note for anyone comparing against our earlier 0.1.31 report (#389): this run holds the vision encoder resident, which costs VRAM and settles the expert slot count at 8,649 instead of ~10,000. Prompt throughput is correspondingly lower; decode is unchanged or slightly higher at 32K. Credentials are stripped from the published config, and `benchmark.py` reads the key from `STRATA_API_KEY`.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.