Pull requests / #993
#993 bench: community report, RTX 5070 Ti + Ryzen 9 5900X, 64 GB DDR4-3600 (IQ2_XS / IQ3_XXS / IQ3_S, six-arm A/B)
closed · @obiscr · 0 comentarios · En GitHub
BenchmarksNVIDIA / CUDAModels & quantsDocumentationWindows
Descripción
Adds `bench/results/2026-10-06-community-rtx5070ti-5900x/` and a row in `docs/COMMUNITY_BENCHMARKS.md`. Single RTX 5070 Ti, Ryzen 9 5900X, 64 GB DDR4-3600, native Windows 10, engine 0.1.39 release build, 65,536-token context. - Three packs (IQ2_XS, IQ3_XXS, IQ3_S), three runs each at 4,096 and 32,768 prompt tokens with the 2026-09-30 community harness; greedy, reasoning off, 256-token cap, a unique nonce per request. Zero reused prompt tokens in every run. - Recall: 9/9 needles at 32k (depths 10/50/90). - Six-arm A/B of the levers from #832 on this box: `--prefill auto:32768` is +29% prefill at 32K; `--kv q4_0` costs 16-18% of prefill and buys nothing; the #832 engine env trio measures zero. Same GPU, different host (Zen 3 / DDR4-3600 / PCIe Gen4 here against Zen 5 / DDR5 / PCIe Gen5 there), so the settings do not transfer. - The README documents the session's decode drift (53.9-63.5 tok/s for one configuration) separately from the prefill numbers, which stayed within ±2%. The second commit fixes the cp936 trap the harness hits on a non-UTF-8 Windows locale (the same class of bug as #881, but here it mis-decodes the tokenizer silently instead of raising UnicodeDecodeError).
En el sitio
Enlaces a install, modelos, releases.