Pull requests / #913
#913 bench: community report, RTX 4060 Ti + Xeon E5-2673 v3 (DDR3, PCIe 3.0 x8), Q2_0 and IQ3_S, 0.1.39 path vs #706 + #764 + a busier tier
closed · @1872183316 · 0 commentaires · Sur GitHub
BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentation
Description
Results-only community benchmark report per `docs/COMMUNITY_BENCHMARKS.md`, in `bench/results/2026-10-05-community-e5-2673v3-rtx-4060ti/` (README, `results.json` with every request, the two configs, the benchmark script and prompts), plus one line in the Community reports list. No engine or setup changes. **PC:** RTX 4060 Ti 16 GB on PCIe 3.0 x8, Xeon E5-2673 v3 (12 cores, AVX2, no AVX-512), 3 x 32 GB DDR3-1866 (64 GB of it interleaved), SATA SSD, Ubuntu 24.04.3. Strata 0.1.39 built from source (CUDA 12.8, sm_89). **Models:** the original GSQ-RCO Q2_0 and IQ3_S (SHA-256 in the README), setup's default config for this PC (64K, int8 KV, `--spec 4 --spec-min-p 0.5`, `--expert-cache auto`). **Arms:** the 0.1.39 code path (`STRATA_Q2_LEGACY=1 STRATA_ADAPT_LAG=1`, default tier) and the same binary with #706, #764 and `--adapt-every 1 --adapt-swaps 160 --adapt-decay 0.97` (#907); alternated, a fresh engine each run, the page cache dropped before each start. 4 short prompts (Chinese essay, code edit, translation, explanation), greedy, 512 tokens, thinking off, 3 rounds x 2 passes. | Model | 0.1.39 path | Improved | | --- | ---: | ---: | | Q2_0 | 48.1 tok/s (48.08 / 48.33 / 48.14) | 55.7 tok/s (55.40 / 55.80 / 55.66) | | IQ3_S | 30.5 tok/s (30.52 / 30.51 / 30.51) | 35.0 tok/s (34.79 / 35.05 / 35.04) | A 3,454-token prompt: Q2_0 715, IQ3_S 439 tok/s in both arms. Per-prompt medians and ranges, hit rates, generated token counts and limits are in the README; the analysis behind the improved arm is in #906. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sur le site
Liens install, modèles, releases.