Pull requests / #913

#913 bench: community report, RTX 4060 Ti + Xeon E5-2673 v3 (DDR3, PCIe 3.0 x8), Q2_0 and IQ3_S, 0.1.39 path vs #706 + #764 + a busier tier

closed · @1872183316 · 0 Kommentare · Auf GitHub

BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentation

Beschreibung

Results-only community benchmark report per `docs/COMMUNITY_BENCHMARKS.md`, in
`bench/results/2026-10-05-community-e5-2673v3-rtx-4060ti/` (README, `results.json` with every request, the two
configs, the benchmark script and prompts), plus one line in the Community reports list. No engine or setup changes.

**PC:** RTX 4060 Ti 16 GB on PCIe 3.0 x8, Xeon E5-2673 v3 (12 cores, AVX2, no AVX-512), 3 x 32 GB DDR3-1866 (64 GB
of it interleaved), SATA SSD, Ubuntu 24.04.3. Strata 0.1.39 built from source (CUDA 12.8, sm_89).

**Models:** the original GSQ-RCO Q2_0 and IQ3_S (SHA-256 in the README), setup's default config for this PC (64K,
int8 KV, `--spec 4 --spec-min-p 0.5`, `--expert-cache auto`).

**Arms:** the 0.1.39 code path (`STRATA_Q2_LEGACY=1 STRATA_ADAPT_LAG=1`, default tier) and the same binary with #706,
#764 and `--adapt-every 1 --adapt-swaps 160 --adapt-decay 0.97` (#907); alternated, a fresh engine each run, the page
cache dropped before each start. 4 short prompts (Chinese essay, code edit, translation, explanation), greedy, 512
tokens, thinking off, 3 rounds x 2 passes.

| Model | 0.1.39 path | Improved |
| --- | ---: | ---: |
| Q2_0 | 48.1 tok/s (48.08 / 48.33 / 48.14) | 55.7 tok/s (55.40 / 55.80 / 55.66) |
| IQ3_S | 30.5 tok/s (30.52 / 30.51 / 30.51) | 35.0 tok/s (34.79 / 35.05 / 35.04) |

A 3,454-token prompt: Q2_0 715, IQ3_S 439 tok/s in both arms. Per-prompt medians and ranges, hit rates, generated
token counts and limits are in the README; the analysis behind the improved arm is in #906.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Mehr auf der Site

Links zu Install, Modellen, Releases.