Pull requests / #538
#538 Community benchmark: RTX 5060 Ti 16 GB + EPYC 7B12, UD-Q4_K_XL at 262,144 tokens
closed · @QilinWan · 0 コメント · GitHub で見る
BenchmarksNVIDIA / CUDAModels & quantsDocumentation
本文
Community benchmark: **RTX 5060 Ti 16 GB + EPYC 7B12**, Strata 0.1.35, the experimental **unsloth UD-Q4_K_XL** weight (111 GB) at a **262,144-token** context, KV `q4_0` with `--kv-resident 32768`, single 16 GB GPU. Prepared per `docs/COMMUNITY_BENCHMARKS.md`. Files: `README.md` (narrative), `matrix.md` / `matrix.json` (tables + machine-readable), `data/` (per-run JSON, engine logs, and the measurement scripts used). Highlights (all from the engine's own timing lines and counters, no estimates): - The whole 24,576-expert weight stays resident in RAM: zero file reads for short prompts, and the decode expert-cache hit rate is 73.1% (8K context). - KV streaming: 95.5%–99.9% of block reads hit VRAM; `32768 of 262144 cells per QSA layer in VRAM`, `K/V in 1.69 GiB of pinned RAM`. - Preferred configuration found locally: the default resident arena (29.45 tok/s) beats a pinned `--resident-budget-gib 71` complement (24.7 tok/s), while `--mmap-experts` is unusable here (7.5 tok/s; one request read 249 GB from the file). - `--spec 2` (window 4) beats the defaults: 26.05 vs 25.35 tok/s and 73.25% vs 62.2% draft acceptance; `--spec 8` collapses to 20.85 tok/s / 40.3%. - Vision: with the encoder on the host CPU (`STRATA_VISION_CUDA=OFF`) the per-image encode is ~0.94 s versus ~0.14 s on the GPU, but it costs no VRAM — 2,644 expert slots versus 2,076 with GPU vision. Both paths read the attached figure correctly. - `--shared-expert-arena` did not reduce startup time here (88 s → 80 s). Build note: this machine has an `sm_120` GPU and needs **CUDA 13.2.86 or newer**; 13.2.78 miscompiles the `IQ3_S`/`IQ2_S` kernels here (the parity tool reports 36 failures at 13.2.78 and 0 at 13.2.86, with `ref rel` going from 0.67–1.03 to 5e-8–6e-5). Mentioned because the report format asks for the CUDA version and changed build options. Run counts are not uniform: the report states per configuration how many runs were measured (some are single runs) and which values could not be measured, rather than filling them in.
関連リンク
インストール・モデル・リリースへの站内リンク。