Issues / #1200

#1200 RTX 5060 Ti 16 GB / Windows / IQ3_S: 36–42 tok/s — tuning directions & method (local gate scan, large pages, PLE FP8)

open · @djaafer1975 · 2 comentários · No GitHub

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsWindows

Descrição

Posting our box's numbers and the tuning **directions/method** as a standalone reference, so it isn't buried inside another thread. Directions only — the exact per-machine values will differ, and that's the point.

**Our machine / model config**
- GPU: RTX 5060 Ti 16 GB (driver 610.88), CPU i5-12600KF, 64 GB system RAM, Windows.
- Model: Qwen3.8-Flash-Next GSQ-RCO **IQ3_S** (~54.8 GB split GGUF) + separate PLE n-gram table + BF16 mmproj for image input.
- Engine profile: KV int8, `--kv-resident 32768`, spec depth 4, `--spec-min-p 0.7`, PLE table on.
- Observed steady-state decode: **36–42 tok/s** on this single 16 GB card (engine-side log numbers).

**Direction 1 — re-scan gate parameters locally; never copy another box's value.**
The default `spec-min-p` 0.5 was not our optimum; a local scan landed on 0.7 (clearly better on Chinese, neutral on code), while other reported setups land at 0.8. Method that made the result trustworthy: two fixed prompt sets (Chinese + code), 2 warm-up runs, median of 3 measured runs per arm, arms = 0.5 / 0.7 / 0.8 / 0.9. Same shape as built-in calibration, but with my own workload instead of synthetic prompts.

**Direction 2 — on Windows, check the memory-page path before tuning anything else.**
Our arena log shows large pages granted (2 MB). If your log says they were refused and it fell back to 4 KB pages across a ~46–50 GB expert pool, that is the first suspect for a 3–5× decode gap — far more likely than GPU or quant choice. Fix: grant the account "Lock pages in memory", then re-run calibrate. Worth comparing `large pages ... (2097152 B)` lines between configs when someone reports unexpectedly low tok/s on a similar card.

**Direction 3 — PLE table precision is a free upgrade if disk allows.**
Swapped the IQ4 n-gram table for the upstream FP8 E4M3 table: byte-for-byte verified against official shards, engine confirms it loads, speed unchanged (36–40 tok/s before and after) — strictly better fidelity at zero latency cost. Caveat: use a mirror that ships real FP8 tensors; the official index.json is missing `weight_scale`.

**Method note (applies to all of the above)**
Read tok/s from engine-side log lines, keep prompt size / max_tokens / temperature fixed across arms, take medians. Wall-clock UI numbers with drifting prompts will happily "prove" a tuning that doesn't exist.

Scene render below is unrelated to performance — a single-file voxel garden (pagoda on a frozen lake at night), just for fun.

![voxel-pagoda](https://gist.githubusercontent.com/djaafer1975/9bd02af0bd81814e6214985c324e3aa9/raw/5486c92eaf8a04123942150fb9554f6c4ba35bda/pagoda-night-snow.jpg)

No site

Links install, modelos, releases.