Pull requests / #357
#357 Start: read the expert arena unbuffered when the file cache cannot help (Windows; part 1 of #285)
closed · @sergqwer · 0 commentaires · Sur GitHub
Setup & installNVIDIA / CUDAModels & quantsWindowsLinux
Description
Part 1 of the #285 split, on 0.1.31. A buffered read copies every expert through the file cache. That pays off only when the cache can keep the files for the next start. On a PC whose RAM cannot hold them beside the arena, every start reads the drive anyway, and the copy is pure cost. For example, a 64 GB PC with IQ2_XS has 24 GiB left for 36.5 GiB of files. What changes: - **`experts_unbuffered()` (pinned.cu) decides the reader.** - Timed random 64 KiB reads tell whether the files are already in the cache: a cached read takes ~10 us, the drive ~80. - Unbuffered is chosen only when the files are not cached **and** the RAM left beside the arena could not keep them either. - So a warm restart (idle unload) stays on 0.1.31's FILE* path. A cold start on a PC with enough RAM also stays buffered, so the restart after it is warm. - `STRATA_UNBUFFERED_LOAD=1/0` forces the choice. The log names the reader and why. - **`load_experts_gguf`, unbuffered:** - each chunk's 4 KiB-aligned window is read with `FILE_FLAG_NO_BUFFERING` into an aligned buffer and scattered into the blobs; - it keeps a handle per role file, as `expert_gguf_file()` resolves it (split shards); - 16 readers. - **`load_experts_direct` (pinned.cu):** a pack's experts.bin ranges are read unbuffered straight into the arena. If the ranges are not 4 KiB-aligned, it falls back to `load_experts_ranges`. - **Linux and macOS read as before:** `experts_unbuffered` returns false there, and both unbuffered readers fall back to the buffered ones. Greedy tokens are identical unbuffered, buffered and on 0.1.31 main. ## Start times Setup: - IQ2_XS without experts.bin, RTX 5090, PCIe 5 drive, `generate` of 16 tokens. - **Cold:** the IQ2_XS files were pushed out of the file cache by reading 142 GB of other files. - **Restart:** the same start again right away. - Each cell shows the time until the arena is loaded / until the 16-token answer is done. **An emulated 64 GB PC** (a locked-pages RAM ballast leaves 58 GiB), large pages, two rounds: | | cold | restart | |---|---|---| | 0.1.31 | 22.1-25.2 / 25.9-29.0 s | 12.5-16.0 / 16.3-19.6 s | | part 1 | 18.3-20.2 / 21.9-23.6 s | **6.0-6.1 / 8.8 s** | Here the probe finds the files uncached on every start, and 58 GiB cannot keep 36.5 GiB of files beside a 33 GiB arena. So part 1 reads unbuffered both times. 0.1.31's restart reads the drive again too: the cache lost the files to the arena. **128 GB, nothing changed**, large pages: | | cold | restart | |---|---|---| | 0.1.31 (3 runs) | 17.6-18.7 / 21.1-22.4 s | 4.6-5.1 / 7.0-7.6 s | | part 1 | 19.3 / 23.1 s | 4.7 / 7.4 s | Both of these starts read through the file cache: - On the cold start the probe found the files uncached (1 of 16 reads). But 84 GiB beside the arena can keep them, so the cache path was chosen, and the restart that followed was warm (15 of 16). - The cold spread is the drive's. Parts 1+2 gave 17.6-19.9 s cold across six runs, with the probe forced off (`STRATA_UNBUFFERED_LOAD=0`) as well as on. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
Sur le site
Liens install, modèles, releases.