Pull requests / #1039
#1039 Start: register the expert arena per layer on a thread ahead of the readers (Windows; part 2 of #285)
closed · @sergqwer · 0 comentarios · En GitHub
Setup & installMulti-GPUNVIDIA / CUDAModels & quantsWindowsLinux
Descripción
Replaces #358. GitHub closed it on 2026-10-05, when my fork was made private by mistake: that took the fork out of the network for good, so it can no longer open pull requests. This is the same branch and commit (`59d97b3`), opened from a new fork; the discussion and the measurements are in #358. --- Part 2 of the #285 split. **One commit on 0.1.38** (part 1, #357, is in 0.1.38). The start times below were measured on 0.1.31; the checksums on 0.1.38 are in the comments. 0.1.31 registers the whole arena with CUDA before the first byte is read. With 4 KB pages, that registration is ~1.5 s of the start that nothing hides (IQ2_XS: 33 GiB). Every account without the Lock Pages privilege gets 4 KB pages. What changes: - **`PinnedArena(bytes, bounds, Deferred)` only reserves the arena.** `register_slices()` registers the per-layer slices in order on a thread and publishes how many are done. - **Every reader writes layer L only once slices L and L+1 are registered.** This covers `load_experts_gguf` (buffered and unbuffered), `load_experts_ranges` and `load_experts_direct`. - Writing a page while `cudaHostRegister` runs on it corrupted the arena under WDDM, even on large pages. - A layer boundary inside a page puts that page in both registrations, so the next slice must be done too. - In effect this gives the page-aligned bounds you asked for, without moving the slice bounds, which an expert must not straddle. - **A slice CUDA refuses ends the registration.** The rest is locked resident, as in the sliced fallback. - **It is used only where the whole arena would be registered anyway:** - no cap (`STRATA_ARENA_PIN_GIB`, multi-GPU under WDDM); - no shared arena file; - Windows only: Linux keeps #253's whole-arena pin unchanged. - **`STRATA_DEFERRED_REGISTER=0`** restores the old order. - **`STRATA_VERIFY_ARENA=1`** prints the loaded arena's FNV-1a checksum (for tests). **Checksums match.** The old order and the deferred one give the same checksum across: - 4 KB pages and large pages; - the GGUF read in place and experts.bin; - buffered and unbuffered reads. ## Start times Setup: - IQ2_XS without experts.bin, RTX 5090, PCIe 5 drive, `generate` of 16 tokens. - **Cold:** the IQ2_XS files were pushed out of the file cache by reading 142 GB of other files. - **Restart:** the same start again right away. - Each cell shows the time until the arena is loaded / until the 16-token answer is done. "Parts 1+2" is this branch. **128 GB** (the file cache can keep the files), two to three runs each: | | cold | restart | |---|---|---| | 4 KB pages (`STRATA_NO_LARGEPAGES=1`), 0.1.31 | 12.4-13.0 / 16.3-16.8 s | 7.8-8.5 / 10.9-11.7 s | | 4 KB pages, parts 1+2 | **11.1-11.5 / 14.9-15.4 s** | **5.5-5.7 / 8.5-8.9 s** | | large pages, 0.1.31 | 17.6-18.7 / 21.1-22.4 s | 4.6-5.1 / 7.0-7.6 s | | large pages, parts 1+2 | 17.6-19.9 / 21.3-23.6 s | 4.1-4.5 / 6.7-7.1 s | The cold spread with large pages comes from the drive: the cache flush reads 142 GB from the same NVMe. A/B runs of this branch with `STRATA_DEFERRED_REGISTER=0` (19.9 s) and with `STRATA_UNBUFFERED_LOAD=0` (17.6 / 18.8 s) fall inside that same spread. **An emulated 64 GB PC** (a locked-pages RAM ballast leaves 58 GiB): - 4 KB pages: two runs each. - Large pages: four runs each. | | cold | restart | |---|---|---| | 4 KB pages, 0.1.31 | 13.0-13.2 / 17.0-17.2 s | 12.9-14.1 / 16.8-18.0 s | | 4 KB pages, parts 1+2 | **7.6-8.2 / 11.1-12.1 s** | **6.0-6.1 / 8.7 s** | | large pages, 0.1.31 | 15.9-21.8 / 19.6-25.5 s | 12.2-22.6 / 15.7-26.2 s | | large pages, parts 1+2 | 12.0-15.5 / 15.8-19.1 s | 6.1-11.9 / 8.7-14.7 s | With large pages under that memory pressure, the allocation itself takes 6-12 s and varies from run to run. Windows has to gather contiguous 2 MB pages before the first byte is read, hence the wide ranges in both rows. ## Part 3 is dropped I built and measured part 3 too: the arena opened on its own thread beside the dense weights, plus a cuBLAS handle made on a thread. It gave no net gain here: - **The arena finished later:** 5.5 vs 4.5 s with large pages, 5.5-5.7 vs 4.9 s with 4 KB pages. The arena thread and cuBLAS init contend with the arena's readers and its registration. - **The first decode step did not move:** it came at 7.1-7.7 s either way. So it is not part of the split. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
En el sitio
Enlaces a install, modelos, releases.