Pull requests / #1751
#1751 bench: community report, RTX 5060 Ti 16 GB (WSL2), GSQ-RCO IQ3_S on engine 0.1.41
open · @kgmkm · 0 comments · View on GitHub
BenchmarksNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Description
## Summary Community benchmark for engine 0.1.41 on one RTX 5060 Ti 16 GB (Ryzen 7 3700X, 128 GB RAM, Windows 11 + WSL2 Ubuntu 24.04), serving the official ISTA-DASLab GSQ-RCO IQ3_S at 131,072 context, single-GPU, with experimental speed projection enabled. Results only: no engine, build or test changes. One run per condition (stated in the README, no median/range), all prompts unique with 0 reused tokens: | Configuration | Prompt tok/s | Decode tok/s | | --- | --- | --- | | decode-story (48 prompt, 1400 gen) | 45.4 | 48.2 | | decode-explan (47 prompt, 1284 gen) | 41.9 | 47.7 | | prompt-4k (3820 prompt) | 343.0 | n/a | | prompt-20k (19196 prompt) | 598.3 | n/a | Draft acceptance 58.7%/61.7%, decode expert cache hit 86.1% best. The README also includes an ESP refusal A/B on the same prompt: ESP on generated 500 tokens (finish=length), ESP off stopped at 106 tokens with a refusal text. ## What changed - Add `bench/results/2026-10-09-community-rtx-5060ti-wsl2-iq3_s-0141/`: the README (hardware, settings, method, tables, limitations), per-condition medians (`measurements.csv`), one row per request (`requests.csv`), a `/metrics` snapshot, the sanitized server config and engine start lines, an engine-log excerpt, and the client script. Local paths are generalized and no credentials are included. - Add nothing else: no code, docs, or index changes (index left to the maintainers). - Not attached: the full multi-GB engine log and the model/pack files. The README lists what is missing. ## Extra Notes - WSL2 + virtiofs stack (WSL components 3.0.1.0, distro in WSL2 mode), so these numbers describe that stack, not native Linux on the same card. - Short-prompt reads carry the 0.1.41 CPU-share caveat noted in the README. - Not tested: needle recall, TTFT, vision, soak, other quants, larger contexts, a layer split across the second card, calibration sweeps.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.