Pull requests / #832
#832 bench: community report, RTX 5070 Ti + Ryzen 7 9800X3D on Windows 11, IQ3_S, 5K to 164K prompt tokens
closed · @Zauberio · 0 comentarios · En GitHub
BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentationWindows
Descripción
Adds `bench/results/2026-10-04-community-rtx-5070ti-windows/`: a community benchmark report for a single RTX 5070 Ti (16 GB) with 64 GB RAM on native Windows 11, engine 0.1.39 release build, IQ3_S native pack, 262K context. - Battery at 5K/41K/131K/164K prompt tokens (greedy, 256-token cap, unique marker per run to defeat the prompt cache). Numbers are the engine's own per-request log lines, not client walls; prefill is stable to ±0.2% run-to-run. - Model files hash-verified against the HF revision's published LFS OIDs (`model-provenance.json`); engine digest verified against the release asset at install. - The harness that issued the requests is included (`battery.ps1`). The README also carries an appendix condensing the A/B work done on this box since 0.1.35: the version ladder (the sm_90+ cluster decode kernels, the 0.1.38 prompt-path rewrite), the env levers kept here (`STRATA_PF_FUSED=1`, `TILE=128`, `GR_DOWN_MAX4=1` — +6.9/+8.1/+2.6% as a trio on 0.1.39, 8-leg ABBA), the calibrated `--pcie-frac 0.35` and prefill cap `auto:32768`, and the measured losers (ADAPT_NOWAIT, IQ_PREFETCH=8192, dma/direct pcie-mode, a stage-overlap patch, router-lookahead prefetch) with their numbers, so nobody re-spends an evening on them. The equal-chunk prefill result measured here (+4.3% at 22.6K tokens) is noted and cross-linked to #693, where the idea lives.
En el sitio
Enlaces a install, modelos, releases.