Pull requests / #202
#202 ple: --ple-io ram locks the n-gram table in RAM (no SSD read on the prompt/token path)
closed · @q8atnight · 0 コメント · GitHub で見る
Setup & installMulti-GPUNVIDIA / CUDAModels & quantsLinux
本文
## Summary A third `--ple-io` mode: `ram` = the `mmap` mapping with the whole table `mlock`ed at open, so no SSD read ever sits on the prompt or token path. `madvise(MADV_WILLNEED)` first, then `mlock`; a start log line reports the time it took. One commit on top of 0.1.27 (a790805). `direct` stays the default and `mmap` is unchanged; the ngram files are untouched upstream, so the change is the `PleIoOptions::lock` flag, its `open()` handling, and the option plumbing. ## Measured On 2x RTX 3090, IQ3_S, 121 GB RAM: stock `--ple-io direct` reads the 28.8 GB table with O_DIRECT, and a cold 32K prompt waited ~3 s on those reads. With `ram` the table is mapped and locked once at start, and no table read touches the SSD again. In use on our box since 2026-09-28 with no issue. ## Cost and failure mode `ram` costs host RAM = the table size (28.8 GB here; our IQ3_S dual setup uses ~82 of 121 GB). If `mlock` is refused — `ulimit -l` too low, Docker's default memlock is 64 KB — the engine does NOT fail: it prints `strata: PLE table mlock failed (...): touching its pages instead` and touches every page, so the table still enters the page cache but the kernel may evict it under memory pressure. For the lock itself to hold, start with `ulimit -l unlimited` (Docker: `--ulimit memlock=-1`). The log line `strata generate: PLE table locked in RAM (--ple-io ram) in X s` distinguishes the two outcomes. ## Switch `--ple-io ram` is opt-in; `direct` (and `mmap`) behave exactly as before.
関連リンク
インストール・モデル・リリースへの站内リンク。