Pull requests / #1751

#1751 bench: community report, RTX 5060 Ti 16 GB (WSL2), GSQ-RCO IQ3_S on engine 0.1.41

open · @kgmkm · 0 comments · View on GitHub

BenchmarksNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

Description

## Summary

Community benchmark for engine 0.1.41 on one RTX 5060 Ti 16 GB
(Ryzen 7 3700X, 128 GB RAM, Windows 11 + WSL2 Ubuntu 24.04), serving the
official ISTA-DASLab GSQ-RCO IQ3_S at 131,072 context, single-GPU, with
experimental speed projection enabled. Results only: no engine, build or
test changes.

One run per condition (stated in the README, no median/range), all prompts
unique with 0 reused tokens:

| Configuration | Prompt tok/s | Decode tok/s |
| --- | --- | --- |
| decode-story (48 prompt, 1400 gen) | 45.4 | 48.2 |
| decode-explan (47 prompt, 1284 gen) | 41.9 | 47.7 |
| prompt-4k (3820 prompt) | 343.0 | n/a |
| prompt-20k (19196 prompt) | 598.3 | n/a |

Draft acceptance 58.7%/61.7%, decode expert cache hit 86.1% best. The README
also includes an ESP refusal A/B on the same prompt: ESP on generated 500
tokens (finish=length), ESP off stopped at 106 tokens with a refusal text.

## What changed

- Add `bench/results/2026-10-09-community-rtx-5060ti-wsl2-iq3_s-0141/`:
  the README (hardware, settings, method, tables, limitations),
  per-condition medians (`measurements.csv`), one row per request
  (`requests.csv`), a `/metrics` snapshot, the sanitized server config and
  engine start lines, an engine-log excerpt, and the client script.
  Local paths are generalized and no credentials are included.
- Add nothing else: no code, docs, or index changes (index left to
  the maintainers).
- Not attached: the full multi-GB engine log and the model/pack files.
  The README lists what is missing.

## Extra Notes

- WSL2 + virtiofs stack (WSL components 3.0.1.0, distro in WSL2 mode),
  so these numbers describe that stack, not native Linux on the same card.
- Short-prompt reads carry the 0.1.41 CPU-share caveat noted in the README.
- Not tested: needle recall, TTFT, vision, soak, other quants, larger
  contexts, a layer split across the second card, calibration sweeps.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.