Pull requests / #1429

#1429 bench: 2x TITAN RTX (sm_75) at Strata v0.1.40.3, IQ3_S at 262K

open · @shrisha108 · 0 评论 · 在 GitHub 查看

BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentationLinux

描述

Community benchmark per `docs/COMMUNITY_BENCHMARKS.md`, re-measured on the current release tag.

**Hardware:** 2x NVIDIA TITAN RTX 24 GB (Turing, sm_75), `nvidia-smi topo -m` reports `NV2` but the engine does not use it, so the same numbers should be expected without a bridge. Xeon E5-2696 v4 (no AVX-512), 125 GiB RAM, PCIe x16 gen 3 under load.

**Software:** AlmaLinux 9.8, driver 615.71.09, CUDA 12.9, Strata tag **`v0.1.40.3`** (commit `d5ea713`) built from source with `CMAKE_CUDA_ARCHITECTURES=75`.

**Configuration:** `--expert-cache auto --prefill auto --spec 4 --mtp ... --max-context 262144 --kv int8 --kv-resident 65536 --vision --vram-reserve-mib 700 --ple-io ram --layer-split auto --reasoning-budget-tokens 12000`.

**Results** (three runs per length, prompt tokens; `benchmark.py` reproduces all nine):

| Prompt | Prompt t/s | Decode t/s | TTFT |
| --- | --- | --- | --- |
| 4,096 | 831.8 / 876.3 / 868.9 | 65.9 / 70.9 / 71.5 | 3.50-3.69 s |
| 32,768 | 1552.1 / 1548.9 / 1540.3 | 77.3 / 79.2 / 77.0 | 15.97-16.10 s |
| 128,000 | 1636.7 / 1626.9 / 1621.7 | 71.1 / 73.8 / 70.7 | 59.13-59.69 s |

8,649 experts cached (16.00 GiB of VRAM), expert cache 95-99% hit. `needle_bench.py --lengths 32k,128k --depths 10,50,90`: **6 of 6 found**, no misses, at 125,918-token actual prompts.

Note for anyone comparing against our earlier 0.1.31 report (#389): this run holds the vision encoder resident, which costs VRAM and settles the expert slot count at 8,649 instead of ~10,000. Prompt throughput is correspondingly lower; decode is unchanged or slightly higher at 32K.

Credentials are stripped from the published config, and `benchmark.py` reads the key from `STRATA_API_KEY`.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。