Issues / #1211
#1211 Bench: --ple-io direct vs. --ple-io ram on system with sufficient headroom of RAM
open · @Thxeverybody · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installModels & quantsDocumentation
描述
Prompted by #746 — that issue establishes `--ple-io ram` as the lever for slow/rotational storage. This reports the complementary case: does it help on fast NVMe + fast DDR5 with plenty of free RAM? Short answer: no. Posting the full series so the null result is on record. **Setup.** RX 9070 XT (16 GB) · Ryzen 7 9700X · 128 GB DDR5-6000 · IQ3_S at 128K context · Strata 0.1.39, model + n-gram table (`--ple-gguf`, 28.8 GB) on NVMe. `mlock` worked with the default `ulimit -l unlimited`; with `ram` the start log shows `PLE table locked in RAM (--ple-io ram) in 5.4 s`, +20 GB resident. **Method.** One identical 9-request series against a freshly started server per variant, switching only the `--ple-io` argument: 3 cold ~78k-token prefills, 4 image+text requests (same image files), 2 long generations, temperature 0. Measured from the per-request `timings`/`usage` the server itself returns; identical prompt/completion token counts confirmed per request for both variants. **Results.** | metric | `--ple-io direct` (default) | `--ple-io ram` | |---|---:|---:| | cold prefill, median of 3 (78k tok) | 1,660 tok/s (range 1,595–1,660) | 1,644 tok/s (range 1,616–2,119) | | decode, median of 9 | 52.6 tok/s (47.3–64.5) | 54.9 tok/s (47.3–64.0) | | image prompts (4x, CPU vision path) | 490–497 tok/s | 489–832 tok/s | | resident RAM cost | — | +20 GB | **Reading.** The two distributions overlap almost completely; per-request variance (±15–25%) dominates any difference between variants. A modern NVMe absorbs the PLE read pattern — 90 B rows, up to 256 in flight — at zero visible cost, so `direct` is genuinely free on this hardware, while `ram` pays the table size in memory for nothing. **Takeaway.** `--ple-io ram` is exactly the right lever for the rotational/slow-drive situation described in #746; on NVMe + fast DDR5 there is no measurable gain and the default is correct. The setup behavior as it stands (opting in when a rotational disk is detected) matches the measurements. If anything: one line in the docs — "no gain expected on NVMe" — could stop people from spending RAM they could use for cache. Happy to re-run anything; the series is a single small script if a repro is useful. This bench was drafted with support of qwen3.8-flash-next-iq3_s under Strata and Hermes
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。