Pull requests / #317
#317 ple: keep the SSD awake while rows are read (STRATA_SSD_KEEPALIVE)
closed · @BlueKingMuch · 0 comments · View on GitHub
NVIDIA / CUDAModels & quantsWindows
Description
Some SSDs drop into a power state after ~250 ms without a command and stall the next reads by 50-150 ms. Here that is a WD_BLACK SN7100 on Windows 11; the NVMe idle timeouts of Windows' power plan do not change it. In decode it hits the first PLE read after a pause: at the start of a request, or after a few rounds whose rows all came from the row cache. `--ple-io ram`, which avoids the reads altogether, is not available on Windows. ## What changes - The reader's I/O worker reads one page of the table (a different one each time) when no read has gone out for 100 ms, until 60 s after the last request for rows; then the SSD may sleep until the next request re-arms it. - A keep-alive read counts when it completes and stays out of the row-read counters and latencies; the `ple io` line adds "SSD kept awake by N reads (slowest X ms)". - A keep-alive read that fails, or cannot be submitted, turns the keep-alive off instead of failing the reader. - `STRATA_SSD_KEEPALIVE=0` turns it off, `=50` shortens the period; `STRATA_SSD_KEEPALIVE_WINDOW` sets the window in seconds. Only with the I/O worker (the default), not with `--ple-sync-submit` or `--ple-io mmap/ram`. - The reader's statistics are now copied and reset under the worker's lock (`snapshot()`), since the worker may be counting a keep-alive read. - The selftest checks the period, the window, the re-arm by a ticket the row cache serves, that rows and read counters are untouched, and that the caller-thread mode has none. The tokens do not change: requests without anything timing-dependent give the same 256 tokens with and without it, greedy and sampled. ## Measured RTX 4080 SUPER 32 GB, Ryzen 7 5800X3D, 64 GB DDR4, WD_BLACK SN7100, Windows 11; IQ3_S, `--max-context 262144`. A local test build of 0.1.30 (30ec18e) with this commit and a few other PRs; the comparison is within that build (`STRATA_SSD_KEEPALIVE=0` against the default). The PLE numbers come from per-request read statistics that build logs (rows, SSD reads, slowest read, time decode waited for rows). | per variant: 2 check requests + 12 requests of 256 tokens, 8K and 64K context | keep-alive off | on (default) | |---|---|---| | requests with a PLE read over 10 ms | 6 of 14 (22, 42, 53, 94, 118, 121 ms) | 0 of 14 | | slowest PLE read | 121 ms | 1.0 ms | | decode waiting for rows, ms per round (mean) | 0.63 | 0.24 | | keep-alive reads per request | - | about 7 | Across the 132 measured requests of the eleven other engine variants in the same session, all with the keep-alive on, none had a PLE read over 10 ms. A scan of the table without the engine, reading it the way the engine does after pauses of every length, gave these stall counts: 0 of 211 bursts after pauses under 200 ms, 35 of 73 after pauses of 200 ms and more, and 0 of 68 of those with the keep-alive on. The cost is one 4 KiB read per 100 ms of otherwise idle SSD time, and only within 60 s of a request. It is on by default for that reason; if you prefer it opt-in, that is the default in one `getenv` line. It touches the same reader as #291 (the FP8 table); the two should combine without trouble, since the keep-alive reads are raw page reads. Claude-Session: https://claude.ai/code/session_01FKZxC8sUotEzdWXws2FTUC
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.