Pull requests / #993

#993 bench: community report, RTX 5070 Ti + Ryzen 9 5900X, 64 GB DDR4-3600 (IQ2_XS / IQ3_XXS / IQ3_S, six-arm A/B)

closed · @obiscr · 0 comentarios · En GitHub

BenchmarksNVIDIA / CUDAModels & quantsDocumentationWindows

Descripción

Adds `bench/results/2026-10-06-community-rtx5070ti-5900x/` and a row in `docs/COMMUNITY_BENCHMARKS.md`.

Single RTX 5070 Ti, Ryzen 9 5900X, 64 GB DDR4-3600, native Windows 10, engine 0.1.39 release
build, 65,536-token context.

- Three packs (IQ2_XS, IQ3_XXS, IQ3_S), three runs each at 4,096 and 32,768 prompt tokens
  with the 2026-09-30 community harness; greedy, reasoning off, 256-token cap, a unique
  nonce per request. Zero reused prompt tokens in every run.
- Recall: 9/9 needles at 32k (depths 10/50/90).
- Six-arm A/B of the levers from #832 on this box: `--prefill auto:32768` is +29% prefill at
  32K; `--kv q4_0` costs 16-18% of prefill and buys nothing; the #832 engine env trio measures
  zero. Same GPU, different host (Zen 3 / DDR4-3600 / PCIe Gen4 here against Zen 5 / DDR5 /
  PCIe Gen5 there), so the settings do not transfer.
- The README documents the session's decode drift (53.9-63.5 tok/s for one configuration)
  separately from the prefill numbers, which stayed within ±2%.

The second commit fixes the cp936 trap the harness hits on a non-UTF-8 Windows locale (the
same class of bug as #881, but here it mis-decodes the tokenizer silently instead of raising
UnicodeDecodeError).

En el sitio

Enlaces a install, modelos, releases.