Pull requests / #538

#538 Community benchmark: RTX 5060 Ti 16 GB + EPYC 7B12, UD-Q4_K_XL at 262,144 tokens

closed · @QilinWan · 0 comentarios · En GitHub

BenchmarksNVIDIA / CUDAModels & quantsDocumentation

Descripción

Community benchmark: **RTX 5060 Ti 16 GB + EPYC 7B12**, Strata 0.1.35, the
experimental **unsloth UD-Q4_K_XL** weight (111 GB) at a **262,144-token** context,
KV `q4_0` with `--kv-resident 32768`, single 16 GB GPU.

Prepared per `docs/COMMUNITY_BENCHMARKS.md`. Files: `README.md` (narrative),
`matrix.md` / `matrix.json` (tables + machine-readable), `data/` (per-run JSON,
engine logs, and the measurement scripts used).

Highlights (all from the engine's own timing lines and counters, no estimates):

- The whole 24,576-expert weight stays resident in RAM: zero file reads for short
  prompts, and the decode expert-cache hit rate is 73.1% (8K context).
- KV streaming: 95.5%–99.9% of block reads hit VRAM; `32768 of 262144 cells per QSA
  layer in VRAM`, `K/V in 1.69 GiB of pinned RAM`.
- Preferred configuration found locally: the default resident arena (29.45 tok/s)
  beats a pinned `--resident-budget-gib 71` complement (24.7 tok/s), while
  `--mmap-experts` is unusable here (7.5 tok/s; one request read 249 GB from the file).
- `--spec 2` (window 4) beats the defaults: 26.05 vs 25.35 tok/s and 73.25% vs 62.2%
  draft acceptance; `--spec 8` collapses to 20.85 tok/s / 40.3%.
- Vision: with the encoder on the host CPU (`STRATA_VISION_CUDA=OFF`) the per-image
  encode is ~0.94 s versus ~0.14 s on the GPU, but it costs no VRAM — 2,644 expert
  slots versus 2,076 with GPU vision. Both paths read the attached figure correctly.
- `--shared-expert-arena` did not reduce startup time here (88 s → 80 s).

Build note: this machine has an `sm_120` GPU and needs **CUDA 13.2.86 or newer**;
13.2.78 miscompiles the `IQ3_S`/`IQ2_S` kernels here (the parity tool reports
36 failures at 13.2.78 and 0 at 13.2.86, with `ref rel` going from 0.67–1.03 to
5e-8–6e-5). Mentioned because the report format asks for the CUDA version and
changed build options.

Run counts are not uniform: the report states per configuration how many runs were
measured (some are single runs) and which values could not be measured, rather than
filling them in.

En el sitio

Enlaces a install, modelos, releases.