Pull requests / #433
#433 bench: community report, RTX 5090 + Ryzen 9 5950X (AVX2), UD-Q4_K_XL / IQ3_S / Swift IQ3_XXS on engines 0.1.31-0.1.39
closed · @brenoperucchi · 0 commentaires · Sur GitHub
BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentationWindows
Description
Community benchmark report under `bench/results/2026-10-01-community-rtx-5090-5950x/`, results only (no engine changes), in the `docs/COMMUNITY_BENCHMARKS.md` format. **Hardware:** RTX 5090 32 GB (PCIe Gen 4 x16, also drives the display), Ryzen 9 5950X (AVX2 only, 16 cores), 96 GB DDR4-3200, NVMe, Windows 11, driver 616.64. **Models:** Unsloth UD-Q4_K_XL (hand-made pack per `docs/UNSLOTH_Q4.md`, 32K context), IQ3_S (65K context) and Swift IQ3_XXS with this PC's production config (32K context), all with `STRATA_IQ_MT_MIN=1` on the IQ sizes. KV int8, MTP on. **Configurations:** engines 0.1.31 to 0.1.39 (release binaries). UD-Q4_K_XL with 72 GiB and 40 GiB RAM budgets, with the 0.1.31 stager values and with `--prefill auto:32768`; on 0.1.38 and 0.1.39 also with `STRATA_UNBUFFERED_LOAD=0`. IQ3_S with `auto` and `auto:32768` on 0.1.34, 0.1.38 and 0.1.39. Swift IQ3_XXS on 0.1.34, 0.1.36 (also with `STRATA_PF_FUSED=1`, speed only), 0.1.37, 0.1.38 and 0.1.39. On 0.1.39, IQ3_S and Swift IQ3_XXS also with the arguments setup 0.1.39 writes for this PC. Prompts built from this repo at `v0.1.32` (103 to 28,886 tokens), 1 warm-up + 3 measured runs each, every run reading its whole prompt fresh. **Findings:** - UD-Q4_K_XL with every expert in memory: 0.1.32 to 0.1.34 read prompts 4-10% slower than 0.1.31; the old stager values bring them within 2%. - UD-Q4_K_XL on 0.1.38: 15-40% slower prompts than on 0.1.34, from the unbuffered file tier (#357/#362); `STRATA_UNBUFFERED_LOAD=0` brings them back. Reported in #577 and fixed in 0.1.39: without the variable, both budgets read within 2% of 0.1.38 with it. - `--prefill auto:32768`: +35% on the 14.7K prompt for UD-Q4_K_XL (0.1.34), +14-31% for IQ3_S at 14.7-28.9K; nothing at 2.7K. On 0.1.39 IQ3_S still gains 14% and 25% at 14.7K and 28.9K. - 0.1.38 reads prompts 11-20% faster on the packs (Swift IQ3_XXS at every size, IQ3_S at 14.7K and above), with the same decode; 0.1.39 is within 1-2% of 0.1.38 there. **Limitations:** one machine, one request at a time, 3 runs per cell, no quality checks (for `STRATA_PF_FUSED=1` see #519), TTFT not measured separately. With 96 GB installed the 40 GiB budget does not reproduce a 64 GB PC's SSD reads.
Sur le site
Liens install, modèles, releases.