Pull requests / #285
#285 Start: read the expert arena unbuffered (the GGUF too, fixes #230), register it while it loads, on its own thread
closed · @sergqwer · 0 Kommentare · Auf GitHub
Setup & installNVIDIA / CUDAModels & quantsWindows
Beschreibung
## Summary
The expert arena's load was the start's long pole. It was a buffered read, then one CUDA registration of the whole
arena; registering 63 GiB of 4 KiB pages alone took 6.6 s. This PR changes four things:
- **Unbuffered reads.**
- A pack's `experts.bin` is read with `FILE_FLAG_NO_BUFFERING` straight into the arena (16 readers, 8 MiB each).
- **The GGUF branch too: this fixes #230.** Setup's default native pack has no `experts.bin`, and its GGUF branch
read through MSVC's `std::ifstream` in 4095-byte pieces: 0.02 GiB/s, 42 minutes in #230. Each chunk's 4 KiB-aligned
window is now read unbuffered into an aligned buffer and scattered into the blobs. Off Windows the old reader
stays.
- **Registration while loading.** The arena is reserved first, and a thread registers its per-layer slices ahead of
the readers (`PinnedArena::Deferred` + `register_slices`). A reader waits for its slice: writing pages while
`cudaHostRegister` runs on them corrupted the arena under WDDM, even on large pages.
- **In parallel.** The arena loads on its own thread while the main thread loads the dense weights, the PLE table,
the MTP drafter and the native head.
- **cuBLAS early.** `gemm_prewarm()` creates the handle and runs its first GEMM off the critical path.
## Measured
Measured on current main (0.1.29, d6708a4) with ISTA-DASLab's IQ2_XS, RTX 5090 32 GB, Ryzen 9 9950X3D, 128 GB DDR5, Windows 11. KL is first-token KL divergence (`STRATA_DUMP_FIRST_LOGITS`, #276); two runs of main itself differ by 0 on 5 of 6 short prompts and by 0.00013 on the sixth. IQ2_XS without `experts.bin` (setup's default pack), a 1K-token prompt and 64 tokens:
| | arena read | start to exit |
| --- | ---: | ---: |
| main (buffered, the file already in the OS cache) | 33 GiB at 6.8 GiB/s | 17.5 s |
| this PR (unbuffered from the NVMe) | 33 GiB in 7.7 s, registration alongside | 11.7 s |
main had the advantage of the OS file cache here (128 GB of RAM). The 0.02 GiB/s of #230 appears on a cold start,
or when RAM is too small to cache the shards; the unbuffered reader does not depend on either. A pack with
`experts.bin` (fork, PCIe 5 drive, 63 GiB): 11.8 GiB/s instead of 3.3, first token after ~8 s instead of ~10.
## Correctness
`STRATA_VERIFY_ARENA=1` prints a checksum of the loaded arena. The unbuffered GGUF loader and the buffered one give
the same checksum (`5217dac097a1751f`), and the generated tokens are unchanged.
## Switch
`STRATA_BUFFERED_LOAD=1` keeps the old loaders.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
Mehr auf der Site
Links zu Install, Modellen, Releases.