Pull requests / #698

#698 Bench: RTX 5090 / 9950X3D (IQ3_S, three-pass A/B) + RTX 2080 Ti / 9900KF (IQ3_XXS, Turing)

closed · @CYoung83 · 0 コメント · GitHub で見る

BenchmarksNVIDIA / CUDAModels & quantsDocumentationWindows

本文

Follow-up to #688, which was closed without comments; the head branch was auto-deleted on close so it cannot be reopened. Resubmitting with a second data point added.

**Package 1 — `bench/results/2026-10-03-dev-5090/`**: RTX 5090 / Ryzen 9 9950X3D, Qwen3.8-Flash-Next IQ3_S. Three labeled passes, 36 cells, all cold-prefill (`reused: 0`), 3 repeats:
- `v0.1.34` (source build, arch 120)
- `v0.1.38` (source build, same args)
- `v0.1.38` with `STRATA_PF_FUSED=1`

Decode medians (tok/s): 1k 176.5 / 168.0 / 190.9, 4k 178.3 / 192.7 / 189.2, 32k 160.9 / 191.7 / 181.3, 128k 143.6 / 152.3 / 164.9. Prefill gains 13-14% at 4k/32k from 0.1.34 to 0.1.38, fused adds 4-11% at 4k and up. Raw per-run draft-acceptance counts are in the JSONs for the #463 question; README makes no decode claims beyond medians because run-to-run spread at fixed config is ±15%.

Harness note: the 128k cells initially 400'd because the harness chars-per-token estimate (3.4) overshoots this corpus (~3.0). Patched harness included: parses the exact prompt token count from the 400 body and rescales (up to 3 retries). Every 128k cell lands under the cap with `reused: 0`.

**Package 2 — `bench/results/2026-10-03-ashley-2080ti/`**: RTX 2080 Ti (sm_75, oldest GPU in the tree so far) / i9-9900KF, 64 GB DDR4-3200, IQ3_XXS, v0.1.38 release build, Windows 11. Measured from a second machine over a 2 ms LAN link; engine/API metrics agree within ~1%. Decode flat 33-38 tok/s across 1k..128k (expert stream riding DDR4 + 2.3 GB auto expert cache, `pcie_frac 0.35` auto-selected), prefill 525 tok/s, 128k TTFT 266 s, 114 MiB VRAM free at 128k int8 KV, no SSD spill. Draft acceptance ~72-75%.

Both packages: README + hw capture + raw per-run JSON + engine log tails + the patched harness used. Single commit per package, only `bench/results/` touched. Happy to reshape, split, or trim to whatever convention you want.

関連リンク

インストール・モデル・リリースへの站内リンク。