Issues / #1211

#1211 Bench: --ple-io direct vs. --ple-io ram on system with sufficient headroom of RAM

open · @Thxeverybody · 0 commentaires · Sur GitHub

BenchmarksSetup & installModels & quantsDocumentation

Description

Prompted by #746 — that issue establishes `--ple-io ram` as the lever for slow/rotational
storage. This reports the complementary case: does it help on fast NVMe + fast DDR5 with
plenty of free RAM? Short answer: no. Posting the full series so the null result is on
record.

**Setup.** RX 9070 XT (16 GB) · Ryzen 7 9700X · 128 GB DDR5-6000 · IQ3_S at 128K context ·
Strata 0.1.39, model + n-gram table (`--ple-gguf`, 28.8 GB) on NVMe. `mlock` worked with
the default `ulimit -l unlimited`; with `ram` the start log shows
`PLE table locked in RAM (--ple-io ram) in 5.4 s`, +20 GB resident.

**Method.** One identical 9-request series against a freshly started server per variant,
switching only the `--ple-io` argument: 3 cold ~78k-token prefills, 4 image+text requests
(same image files), 2 long generations, temperature 0. Measured from the per-request
`timings`/`usage` the server itself returns; identical prompt/completion token counts
confirmed per request for both variants.

**Results.**

| metric | `--ple-io direct` (default) | `--ple-io ram` |
|---|---:|---:|
| cold prefill, median of 3 (78k tok) | 1,660 tok/s (range 1,595–1,660) | 1,644 tok/s (range 1,616–2,119) |
| decode, median of 9 | 52.6 tok/s (47.3–64.5) | 54.9 tok/s (47.3–64.0) |
| image prompts (4x, CPU vision path) | 490–497 tok/s | 489–832 tok/s |
| resident RAM cost | — | +20 GB |

**Reading.** The two distributions overlap almost completely; per-request variance (±15–25%)
dominates any difference between variants. A modern NVMe absorbs the PLE read pattern —
90 B rows, up to 256 in flight — at zero visible cost, so `direct` is genuinely free on this
hardware, while `ram` pays the table size in memory for nothing.

**Takeaway.** `--ple-io ram` is exactly the right lever for the rotational/slow-drive
situation described in #746; on NVMe + fast DDR5 there is no measurable gain and the default
is correct. The setup behavior as it stands (opting in when a rotational disk is detected)
matches the measurements. If anything: one line in the docs — "no gain expected on NVMe" —
could stop people from spending RAM they could use for cache.

Happy to re-run anything; the series is a single small script if a repro is useful.

This bench was drafted with support of qwen3.8-flash-next-iq3_s under Strata and Hermes

Sur le site

Liens install, modèles, releases.