Pull requests / #699
#699 Linux: read ahead at startup (cold start ~920 s → 70 s)
closed · @jesdga95 · 0 コメント · GitHub で見る
BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux
本文
On Linux every startup read went to the drive one ~128 KB request at a time (page faults through the mmap, or plain freads with the default readahead), so a cold start sat at queue depth ~2 for about fifteen minutes. Most of it was "filling the GPU's expert cache" at 30-39 MB/s from an NVMe that does ~3 GB/s with a deep queue. Windows already has an unbuffered, overlapped read path for this (`set_unbuffered` / `read_direct`); on Linux `set_unbuffered` returns false, so these reads went through the page cache one fault at a time. This asks the kernel for the pages ahead of each read (madvise / posix_fadvise WILLNEED, in 128 KiB steps because WILLNEED only reads one readahead window per call) for the weights, the native dense matrices, the profile fill (256 pairs ahead), the peer tier's fill (`--peer-device`), the resident RAM copy, the MTP draft files, the native output head and token embedding, and strata-vision's projector. `STRATA_READ_AHEAD=0` turns it all off, and each phase now logs how long it took. Ryzen 9 9950X, RTX 5090, 32 GB DDR5, Gen3 NVMe, Ubuntu (kernel 7.0), v0.1.38, cold page cache: - Q2_0 resident, 262K: ready in **70 s, was ~920 s**. The fill went from ~39 MB/s to 3.2 GB/s (21.47 GiB in 7.1 s). - IQ3_XXS resident + vision + peer tier (5060 Ti): ready in **32 s**. MTP draft files 47 s → 0.3 s, peer fill 13.5 GiB in 5 s, 21.9 GiB RAM copy in 8 s, head/embedding/projector under a second each. - Decode unchanged: struct still gives the same 7,230 tokens, and the bench is within noise. Only tested on Linux with CUDA. On Windows the new helper is a no-op (the unbuffered path covers it), and I haven't compiled it on Windows or HIP. With a layer split, the first card's fill gets the read-ahead (q8atnight measured it below); the other cards' fills don't yet.
関連リンク
インストール・モデル・リリースへの站内リンク。