Pull requests / #202

#202 ple: --ple-io ram locks the n-gram table in RAM (no SSD read on the prompt/token path)

closed · @q8atnight · 0 コメント · GitHub で見る

Setup & installMulti-GPUNVIDIA / CUDAModels & quantsLinux

本文

## Summary

A third `--ple-io` mode: `ram` = the `mmap` mapping with the whole table `mlock`ed at open, so no SSD read ever
sits on the prompt or token path. `madvise(MADV_WILLNEED)` first, then `mlock`; a start log line reports the time
it took. One commit on top of 0.1.27 (a790805). `direct` stays the default and `mmap` is unchanged; the ngram
files are untouched upstream, so the change is the `PleIoOptions::lock` flag, its `open()` handling, and the
option plumbing.

## Measured

On 2x RTX 3090, IQ3_S, 121 GB RAM: stock `--ple-io direct` reads the 28.8 GB table with O_DIRECT, and a cold 32K
prompt waited ~3 s on those reads. With `ram` the table is mapped and locked once at start, and no table read
touches the SSD again. In use on our box since 2026-09-28 with no issue.

## Cost and failure mode

`ram` costs host RAM = the table size (28.8 GB here; our IQ3_S dual setup uses ~82 of 121 GB). If `mlock` is
refused — `ulimit -l` too low, Docker's default memlock is 64 KB — the engine does NOT fail: it prints
`strata: PLE table mlock failed (...): touching its pages instead` and touches every page, so the table still
enters the page cache but the kernel may evict it under memory pressure. For the lock itself to hold, start with
`ulimit -l unlimited` (Docker: `--ulimit memlock=-1`). The log line
`strata generate: PLE table locked in RAM (--ple-io ram) in X s` distinguishes the two outcomes.

## Switch

`--ple-io ram` is opt-in; `direct` (and `mmap`) behave exactly as before.

関連リンク

インストール・モデル・リリースへの站内リンク。