Pull requests / #1016

#1016 Bench: RTX 5060 Ti 16 GB on Ryzen 9 7945HX, IQ3_S

closed · @avarakin · 0 コメント · GitHub で見る

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentationLinux

本文

Results-only report following `docs/COMMUNITY_BENCHMARKS.md`. No engine changes.

## Hardware

- NVIDIA GeForce RTX 5060 Ti, 16,311 MiB VRAM, 180 W power limit, PCIe Gen 5 x8, clocks not fixed.
- AMD Ryzen 9 7945HX (32 logical CPUs); the engine used 15 expert-pool workers plus its host thread.
- `MemTotal` 60.53 GiB (`lsmem` 65.5 GiB online; installed DIMM size not measured), 31.99 GiB swap.
- Model, pack and MTP files on ext4 on a Samsung 970 EVO Plus 2 TB NVMe, reached through a mergerfs union at `/data` whose second branch is ZFS on a 9.1 TB HGST HDD; every Strata file resolves to the NVMe branch.

## Software

Arch Linux, kernel 7.2.8-arch1-1, NVIDIA driver 615.71.09, NVCC 13.4.92, Python 3.14.7. Source build at `d9ab8435f654c368c586340d490915f6addf56a3`, engine 0.1.35, CUDA arch 120, GPU vision helper built.

## Model and configuration

`ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF` **IQ3_S** (the existing community report used IQ2_XS), native IQ3_S pack (1.5 GB) and Q2_0 MTP draft pack (787 MB) prepared by the setup script. Context 131,072; KV `q4_0`, no resident KV; expert cache `auto` = **3,652 slots / 6.98 GiB**, `PROFILE` policy, pre-filled, no eviction; prefill `auto` = 8,192-token chunks borrowing 2,331 cache slots; `--spec 4 --spec-min-p 0.70`; `--pcie-frac 0.00`; low-RAM off; no calibration, no experimental speed projection; reasoning off, temperature 0, 256-token output cap.

## What was measured

The same script and prompts as the RTX 5090 report: one warm-up (excluded) plus three serial runs each at 4,096 / 32,768 / 128,000 prompt tokens, all with **zero reused tokens**, plus six `tools/needle_bench.py` recall checks (all six found).

| Prompt tokens | Prompt tok/s | Decode tok/s | TTFT s | Total s |
| ---: | ---: | ---: | ---: | ---: |
| 4,096 | 1,157.7 [1,145.7–1,157.8] | 56.6 [47.5–58.5] | 3.573 | 8.071 |
| 32,768 | 1,225.3 [1,223.8–1,227.4] | 60.2 [58.9–60.3] | 26.811 | 31.031 |
| 128,000 | 1,175.9 [1,175.5–1,175.9] | 57.2 [55.6–59.7] | 109.024 | 113.471 |

Median [min–max] of three runs. GPU memory peak 15,702 MiB; host RAM peak 55.26 GiB (`MemTotal - MemAvailable`), sampled only during the recall phase.

## Limitations

- **The benchmarking agent runs on the server it measured.** The engine answers one request at a time, and scanning every `prompt … tokens` line between the first 4,096 run and the last 128,000 run finds only the nine benchmark lines (same across the six recall runs), so the measured requests did not overlap other traffic. The only nearby traffic was the agent's own three turns of 27,735 / 27,890 / 28,074 tokens, which ended about one second before the warm-up. The measurements therefore started straight after other work rather than during it.
- The first 4K run (47.5 tok/s decode, 8.97 s total) is the low outlier and followed that 301-token turn directly; no cause is proven. Runs 2 and 3 were 58.5 and 56.6 tok/s.
- Not an isolated machine: other CPU services running, GPU clocks unpinned, expert cache pre-filled at startup and already used by earlier traffic, page cache warm, 548 requests served before the run.
- Measured decode expert-cache hit rate was **44.9–65.4%** with 3,652 slots, against the 97.8–99.7% reported for the RTX 5090's 17,463 slots. Together with the weaker GPU this is the likely reason for the decode gap, but quantization, KV format, engine version, CPU, RAM and PCIe width all differ too, so the two tables are **not** a GPU-only comparison.
- Model GGUF hashes and the Hugging Face revision were not re-verified; the files on disk are dated 2026-10-02.
- Long output, sampled decoding, thinking, vision, tool use, concurrency and sustained thermal behaviour were not evaluated.

## Contents

`README.md` plus per-run request bodies and full streamed chunk logs, `results.json`, `summary.json`, `needles.json`, `telemetry.jsonl`, `memory-summary.json`, the live engine config, `BUILD.json`, and an engine log excerpt. `benchmark.py` here is the community script plus one fix: it matches the engine record to its own request instead of taking `metrics["requests"][0]`, which matters when the server is shared (the copy in the RTX 5090 folder is left untouched). Its `concurrent_requests` counter also counts the script's own warm-up, so it is an upper bound rather than proof of another client.

関連リンク

インストール・モデル・リリースへの站内リンク。