Pull requests / #317

#317 ple: keep the SSD awake while rows are read (STRATA_SSD_KEEPALIVE)

closed · @BlueKingMuch · 0 commentaires · Sur GitHub

NVIDIA / CUDAModels & quantsWindows

Description

Some SSDs drop into a power state after ~250 ms without a command and stall the next reads by 50-150 ms. Here that is a WD_BLACK SN7100 on Windows 11; the NVMe idle timeouts of Windows' power plan do not change it. In decode it hits the first PLE read after a pause: at the start of a request, or after a few rounds whose rows all came from the row cache. `--ple-io ram`, which avoids the reads altogether, is not available on Windows.

## What changes

- The reader's I/O worker reads one page of the table (a different one each time) when no read has gone out for 100 ms, until 60 s after the last request for rows; then the SSD may sleep until the next request re-arms it.
- A keep-alive read counts when it completes and stays out of the row-read counters and latencies; the `ple io` line adds "SSD kept awake by N reads (slowest X ms)".
- A keep-alive read that fails, or cannot be submitted, turns the keep-alive off instead of failing the reader.
- `STRATA_SSD_KEEPALIVE=0` turns it off, `=50` shortens the period; `STRATA_SSD_KEEPALIVE_WINDOW` sets the window in seconds. Only with the I/O worker (the default), not with `--ple-sync-submit` or `--ple-io mmap/ram`.
- The reader's statistics are now copied and reset under the worker's lock (`snapshot()`), since the worker may be counting a keep-alive read.
- The selftest checks the period, the window, the re-arm by a ticket the row cache serves, that rows and read counters are untouched, and that the caller-thread mode has none.

The tokens do not change: requests without anything timing-dependent give the same 256 tokens with and without it, greedy and sampled.

## Measured

RTX 4080 SUPER 32 GB, Ryzen 7 5800X3D, 64 GB DDR4, WD_BLACK SN7100, Windows 11; IQ3_S, `--max-context 262144`. A local test build of 0.1.30 (30ec18e) with this commit and a few other PRs; the comparison is within that build (`STRATA_SSD_KEEPALIVE=0` against the default). The PLE numbers come from per-request read statistics that build logs (rows, SSD reads, slowest read, time decode waited for rows).

| per variant: 2 check requests + 12 requests of 256 tokens, 8K and 64K context | keep-alive off | on (default) |
|---|---|---|
| requests with a PLE read over 10 ms | 6 of 14 (22, 42, 53, 94, 118, 121 ms) | 0 of 14 |
| slowest PLE read | 121 ms | 1.0 ms |
| decode waiting for rows, ms per round (mean) | 0.63 | 0.24 |
| keep-alive reads per request | - | about 7 |

Across the 132 measured requests of the eleven other engine variants in the same session, all with the keep-alive on, none had a PLE read over 10 ms.

A scan of the table without the engine, reading it the way the engine does after pauses of every length, gave these stall counts: 0 of 211 bursts after pauses under 200 ms, 35 of 73 after pauses of 200 ms and more, and 0 of 68 of those with the keep-alive on.

The cost is one 4 KiB read per 100 ms of otherwise idle SSD time, and only within 60 s of a request. It is on by default for that reason; if you prefer it opt-in, that is the default in one `getenv` line. It touches the same reader as #291 (the FP8 table); the two should combine without trouble, since the keep-alive reads are raw page reads.

Claude-Session: https://claude.ai/code/session_01FKZxC8sUotEzdWXws2FTUC

Sur le site

Liens install, modèles, releases.