Pull requests / #362
#362 File tier: read the GGUF in place unbuffered when the file cache cannot keep it beside the RAM budget (Windows; #286 ported into the existing tier)
closed · @sergqwer · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installNVIDIA / CUDAModels & quantsWindowsLinux
描述
Replaces #286, as you asked: whatever won goes into 0.1.31's tier, with no second tier. **Stacked on #357** because it uses its `experts_unbuffered()`. Please review only the last commit. ## What came over from #286 - **Unbuffered reads of the experts outside the RAM budget: ported.** They now read the GGUF in place, one 4 KiB-aligned window per role, instead of experts.bin. This covers the prompt path's stager (`copy_blob`), the decode misses (`prefetch`), the complement copy and the profile fill. Each batch has all its windows in flight at once, overlapped. - **The pinned budget: nothing to port.** 0.1.31's `cudaHostAlloc` complement is page-locked already. It only shrank because the mapped profile fill had taken the RAM first (16 GiB asked, 11.5-13.6 GiB held). With unbuffered reads it keeps all 16 GiB. - **The VRAM <-> RAM exchange: nothing to port.** `stage_exchange` / `commit_exchanges` already swap without a file read. - **The routing prefetch is off when unbuffered** (`warms()` returns false). Its `PrefetchVirtualMemory` of the mapping would read the experts through the file cache a second time. Here it had warmed only 8% of the file tier's decode reads. The choice is `experts_unbuffered(files, budget)`. It goes unbuffered whenever the files could not be kept beside the budget, cached now or not. A partly cached file's mapped pages still land in the working set: on an emulated 64 GB PC, with 14 of 16 probe reads cached, a mapped profile fill shrank the budget to 31 GiB and then ran the RAM out. A 64 GB PC with the same budget, the plain `--mmap-experts` mode without a budget, and Linux keep the mapped path unchanged. `STRATA_UNBUFFERED_LOAD=1/0` forces the choice. ## NTFS and mapped files NTFS runs a file's unbuffered reads one at a time while the file is mapped (or cached) anywhere, and the engine always maps the GGUF shards for the token embedding and PLE. Measured on this PCIe 5 drive, random reads, 6 threads x queue depth 48: | | 512 KiB reads | 8 MiB reads | |---|---:|---:| | the file mapped by nobody | 10.4 GB/s | | | the file mapped (a view or just the section) or read through the cache | 3.4 GB/s | 8.6 GB/s | So windows less than 1 MiB apart are merged into requests of up to 32 MiB. That took the complement copy from 7.4 to 4.5 s. Scattered reads (profile fill, prompt, decode) stay at the one-at-a-time rate. ## Measured Setup, the same as in my #286 reply: - An emulated 32 GB PC with a 16 GB GPU: a locked-pages ballast leaves 26 GiB, and `--vram-reserve-mib 16900`. - IQ2_XS, `--resident-budget-gib 16`, the GGUF in place. - The 8K prompt, then 256 tokens. - Two runs each, all on the same day. | | RAM budget held | decode | answer done (from launch) | prompt: host streaming | engine's TTFT | |---|---:|---:|---:|---:|---:| | 0.1.31 | 11.5 / 13.6 GiB | 69.2 / 66.7 tok/s | 47.3 / 45.1 s | 11.6 / 10.8 s | 27.0 / 28.8 s | | **this** | **16.0 / 16.0 GiB** | **84.3 / 88.9 tok/s** | **21.5 / 22.1 s** | **4.4 / 4.3 s** | **12.6 / 13.4 s** | | old #286 tier (experts.bin) | 16.0 / 16.0 GiB | 90.1 / 87.2 tok/s | 16.7 / 14.6 s | 3.3 / 2.9 s | 6.4 / 5.4 s | Start-up phases: - **Profile fill:** 3.6 s instead of 10.3-14.3 s. It reads batches of 64, the next batch while the current one is copied. - **Complement:** 4.5 s for 16 GiB instead of 9.3-11.4 s for 11.5-13.6 GiB. **Update: a pack's experts.bin too** (second commit). setup's low-RAM packs have an experts.bin, and 0.1.31's tier mapped it. Yesterday that ran the emulated 32 GB PC out of RAM, below 1 GB available. Now it is read unbuffered as well: one window per blob, merged where near. | same setup, experts.bin pack | decode | answer done | prompt: host streaming | |---|---:|---:|---:| | this, experts.bin | **100.2 / 100.6 tok/s** | **17.2 / 15.5 s** | **2.1 s** | | this, the GGUF in place (table above) | 84.3 / 88.9 tok/s | 21.5 / 22.1 s | 4.4 / 4.3 s | Greedy tokens are identical unbuffered, through the cache and on 0.1.31, with either pack. Most of the remaining gap to the old tier is the mapping effect above: the old tier read experts.bin, which nothing mapped. Two possible follow-ups: - read the profile fill's and the complement's experts in one sequential pass per layer; - have the tier read an experts.bin when the pack has one. Greedy tokens are identical unbuffered (forced on a 128 GB run), through the cache, and on 0.1.31. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。