Pull requests / #538
#538 Community benchmark: RTX 5060 Ti 16 GB + EPYC 7B12, UD-Q4_K_XL at 262,144 tokens
closed · @QilinWan · 0 Kommentare · Auf GitHub
BenchmarksNVIDIA / CUDAModels & quantsDocumentation
Beschreibung
Community benchmark: **RTX 5060 Ti 16 GB + EPYC 7B12**, Strata 0.1.35, the experimental **unsloth UD-Q4_K_XL** weight (111 GB) at a **262,144-token** context, KV `q4_0` with `--kv-resident 32768`, single 16 GB GPU. Prepared per `docs/COMMUNITY_BENCHMARKS.md`. Files: `README.md` (narrative), `matrix.md` / `matrix.json` (tables + machine-readable), `data/` (per-run JSON, engine logs, and the measurement scripts used). Highlights (all from the engine's own timing lines and counters, no estimates): - The whole 24,576-expert weight stays resident in RAM: zero file reads for short prompts, and the decode expert-cache hit rate is 73.1% (8K context). - KV streaming: 95.5%–99.9% of block reads hit VRAM; `32768 of 262144 cells per QSA layer in VRAM`, `K/V in 1.69 GiB of pinned RAM`. - Preferred configuration found locally: the default resident arena (29.45 tok/s) beats a pinned `--resident-budget-gib 71` complement (24.7 tok/s), while `--mmap-experts` is unusable here (7.5 tok/s; one request read 249 GB from the file). - `--spec 2` (window 4) beats the defaults: 26.05 vs 25.35 tok/s and 73.25% vs 62.2% draft acceptance; `--spec 8` collapses to 20.85 tok/s / 40.3%. - Vision: with the encoder on the host CPU (`STRATA_VISION_CUDA=OFF`) the per-image encode is ~0.94 s versus ~0.14 s on the GPU, but it costs no VRAM — 2,644 expert slots versus 2,076 with GPU vision. Both paths read the attached figure correctly. - `--shared-expert-arena` did not reduce startup time here (88 s → 80 s). Build note: this machine has an `sm_120` GPU and needs **CUDA 13.2.86 or newer**; 13.2.78 miscompiles the `IQ3_S`/`IQ2_S` kernels here (the parity tool reports 36 failures at 13.2.78 and 0 at 13.2.86, with `ref rel` going from 0.67–1.03 to 5e-8–6e-5). Mentioned because the report format asks for the CUDA version and changed build options. Run counts are not uniform: the report states per configuration how many runs were measured (some are single runs) and which values could not be measured, rather than filling them in.
Mehr auf der Site
Links zu Install, Modellen, Releases.