Pull requests / #538

#538 Community benchmark: RTX 5060 Ti 16 GB + EPYC 7B12, UD-Q4_K_XL at 262,144 tokens

closed · @QilinWan · 0 Kommentare · Auf GitHub

BenchmarksNVIDIA / CUDAModels & quantsDocumentation

Beschreibung

Community benchmark: **RTX 5060 Ti 16 GB + EPYC 7B12**, Strata 0.1.35, the
experimental **unsloth UD-Q4_K_XL** weight (111 GB) at a **262,144-token** context,
KV `q4_0` with `--kv-resident 32768`, single 16 GB GPU.

Prepared per `docs/COMMUNITY_BENCHMARKS.md`. Files: `README.md` (narrative),
`matrix.md` / `matrix.json` (tables + machine-readable), `data/` (per-run JSON,
engine logs, and the measurement scripts used).

Highlights (all from the engine's own timing lines and counters, no estimates):

- The whole 24,576-expert weight stays resident in RAM: zero file reads for short
  prompts, and the decode expert-cache hit rate is 73.1% (8K context).
- KV streaming: 95.5%–99.9% of block reads hit VRAM; `32768 of 262144 cells per QSA
  layer in VRAM`, `K/V in 1.69 GiB of pinned RAM`.
- Preferred configuration found locally: the default resident arena (29.45 tok/s)
  beats a pinned `--resident-budget-gib 71` complement (24.7 tok/s), while
  `--mmap-experts` is unusable here (7.5 tok/s; one request read 249 GB from the file).
- `--spec 2` (window 4) beats the defaults: 26.05 vs 25.35 tok/s and 73.25% vs 62.2%
  draft acceptance; `--spec 8` collapses to 20.85 tok/s / 40.3%.
- Vision: with the encoder on the host CPU (`STRATA_VISION_CUDA=OFF`) the per-image
  encode is ~0.94 s versus ~0.14 s on the GPU, but it costs no VRAM — 2,644 expert
  slots versus 2,076 with GPU vision. Both paths read the attached figure correctly.
- `--shared-expert-arena` did not reduce startup time here (88 s → 80 s).

Build note: this machine has an `sm_120` GPU and needs **CUDA 13.2.86 or newer**;
13.2.78 miscompiles the `IQ3_S`/`IQ2_S` kernels here (the parity tool reports
36 failures at 13.2.78 and 0 at 13.2.86, with `ref rel` going from 0.67–1.03 to
5e-8–6e-5). Mentioned because the report format asks for the CUDA version and
changed build options.

Run counts are not uniform: the report states per configuration how many runs were
measured (some are single runs) and which values could not be measured, rather than
filling them in.

Mehr auf der Site

Links zu Install, Modellen, Releases.