Pull requests / #417
#417 bench: community report, RTX PRO 4500 x1/x2 + RTX PRO 4000, Threadripper PRO 3975WX, engine 0.1.31
closed · @mlfather · 0 Kommentare · Auf GitHub
BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsDocumentation
Beschreibung
# Community benchmark: RTX PRO 4500 Blackwell (x1, x2) + RTX PRO 4000, Threadripper PRO 3975WX, 512 GB RAM Adds `bench/results/2026-10-01-community-rtx-pro-4500/` (README with method, per-pack tables, split decisions, needle recall, limitations; per-run JSON; server configs; engine logs; per-card nvidia-smi samples; BUILD.json). Extends the report posted in #402 with two- and three-GPU layer-split runs, as requested there. **Setup.** Engine 0.1.31, source build, commit `9259cad`, CUDA 13.3, sm_120. Host: Threadripper PRO 3975WX (AVX2, no AVX-512), 512 GB DDR4-8ch, NVMe. Packs on the same binary: Unsloth **UD-Q4_K_XL** (imported per `docs/UNSLOTH_Q4.md`, GGUF read in place, no `experts.bin`), Unsloth **UD-IQ4_XS**, ISTA-DASLab **IQ3_S**. Prompts: 2,700-token review, 370-token codegen, and the same padded to 63,997 and 127,999 tokens; greedy, thinking off, 1 warm-up + 3 (short/codegen) or 2 (long) runs; context limits 8,192 / 81,920 / 139,264, one server per limit. Single card: `--resident-budget-gib 100` (every expert cached, no file reads). Two and three cards: `--layer-split auto` with the default cache mode (the resident budget cannot be combined with a split). **Decode tok/s, median, 1 / 2 / 3 GPUs** (1 = one RTX PRO 4500; 2 = two 4500; 3 = two 4500 + RTX PRO 4000): | Prompt | UD-Q4_K_XL | UD-IQ4_XS | IQ3_S | |---|---|---|---| | short 2.7K | 61.5 / 111.5 / 112.6 | 82.0 / 121.7 / 112.0 | 106.3 / 130.2 / 131.7 | | codegen 370 | 72.4 / 142.4 / 137.9 | 98.0 / 151.9 / 145.6 | 123.8 / 157.3 / 151.6 | | long 64K | 72.1 / 134.8 / 135.8 | 95.8 / 149.2 / 145.7 | 120.3 / 159.1 / 150.8 | | long 128K | 70.2 / 130.2 / 131.5 | 91.4 / 147.8 / n/a | 116.5 / 153.5 / 144.5 | **Prefill tok/s and time to first token, 1 / 2 / 3 GPUs:** | Prompt | UD-Q4_K_XL | IQ3_S | |---|---|---| | long 64K | 2,454 (26.1 s) / 4,920 (13.0 s) / 4,265 (15.0 s) | 3,281 (19.5 s) / 5,643 (11.3 s) / 2,340 (27.4 s) | | long 128K | 2,351 (54.4 s) / 5,124 (25.0 s) / 4,093 (31.3 s) | 3,102 (41.3 s) / 5,757 (22.2 s) / 2,295 (55.8 s) | | long 256K (2 GPUs only) | 4,884 (52.4 s), decode 114.7 tok/s, cache hit 98 % | not run | Peak VRAM 31.2-31.4 GB per card at every setting; host RAM 101-104 GiB (UD-Q4_K_XL, single card with the resident budget), 74-78 GiB (IQ3_S). GPU power 140-195 W. **Observations.** (0) The model's full 262K window is usable on two cards: a 256,011-token prompt reads at 4,884 tok/s (52 s to first token) and decodes at 115 tok/s, 12 % under the 128K figure as the KV cache takes room from the expert cache. (1) With two cards the caches hold ~99.5 % of the routed expert mass (UD-Q4_K_XL K=23 split, 16.6K of 24.6K pairs), which is why the Q4_K pack nearly doubles: its experts otherwise run on ggml's per-token AVX2 path. (2) A third card adds nothing to decode and slows long-prompt prefill (two card boundaries, the slowest card last). (3) `STRATA_KQ256=1` on UD-Q4_K_XL: within noise (67.6 vs 61.5 short, 71.7 vs 72.4 codegen). (4) Needle recall at 32K/128K, depths 10/50/90: 6/6 found. (5) UD-IQ4_XS imports on the unpatched 0.1.31 (on 0.1.18 it needed a local patch). (6) `docs/UNSLOTH_Q4.md`'s serve example omits `--serve`; the binary exits asking for `--tokens` without it. Limitations: single request at a time, greedy only, two runs for long prompts, multi-GPU runs in the default cache mode, no quality suite beyond the needle check. Results-only submission; no engine changes. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01V3YDbe7CsrJt7tTWsck1Uq
Mehr auf der Site
Links zu Install, Modellen, Releases.