Issues / #1200
#1200 RTX 5060 Ti 16 GB / Windows / IQ3_S: 36–42 tok/s — tuning directions & method (local gate scan, large pages, PLE FP8)
open · @djaafer1975 · 2 comentarios · En GitHub
BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsWindows
Descripción
Posting our box's numbers and the tuning **directions/method** as a standalone reference, so it isn't buried inside another thread. Directions only — the exact per-machine values will differ, and that's the point. **Our machine / model config** - GPU: RTX 5060 Ti 16 GB (driver 610.88), CPU i5-12600KF, 64 GB system RAM, Windows. - Model: Qwen3.8-Flash-Next GSQ-RCO **IQ3_S** (~54.8 GB split GGUF) + separate PLE n-gram table + BF16 mmproj for image input. - Engine profile: KV int8, `--kv-resident 32768`, spec depth 4, `--spec-min-p 0.7`, PLE table on. - Observed steady-state decode: **36–42 tok/s** on this single 16 GB card (engine-side log numbers). **Direction 1 — re-scan gate parameters locally; never copy another box's value.** The default `spec-min-p` 0.5 was not our optimum; a local scan landed on 0.7 (clearly better on Chinese, neutral on code), while other reported setups land at 0.8. Method that made the result trustworthy: two fixed prompt sets (Chinese + code), 2 warm-up runs, median of 3 measured runs per arm, arms = 0.5 / 0.7 / 0.8 / 0.9. Same shape as built-in calibration, but with my own workload instead of synthetic prompts. **Direction 2 — on Windows, check the memory-page path before tuning anything else.** Our arena log shows large pages granted (2 MB). If your log says they were refused and it fell back to 4 KB pages across a ~46–50 GB expert pool, that is the first suspect for a 3–5× decode gap — far more likely than GPU or quant choice. Fix: grant the account "Lock pages in memory", then re-run calibrate. Worth comparing `large pages ... (2097152 B)` lines between configs when someone reports unexpectedly low tok/s on a similar card. **Direction 3 — PLE table precision is a free upgrade if disk allows.** Swapped the IQ4 n-gram table for the upstream FP8 E4M3 table: byte-for-byte verified against official shards, engine confirms it loads, speed unchanged (36–40 tok/s before and after) — strictly better fidelity at zero latency cost. Caveat: use a mirror that ships real FP8 tensors; the official index.json is missing `weight_scale`. **Method note (applies to all of the above)** Read tok/s from engine-side log lines, keep prompt size / max_tokens / temperature fixed across arms, take medians. Wall-clock UI numbers with drifting prompts will happily "prove" a tuning that doesn't exist. Scene render below is unrelated to performance — a single-file voxel garden (pagoda on a frozen lake at night), just for fun. 
En el sitio
Enlaces a install, modelos, releases.