Pull requests / #358

#358 Start: register the expert arena per layer on a thread ahead of the readers (Windows; part 2 of #285)

closed · @sergqwer · 0 commentaires · Sur GitHub

Setup & installMulti-GPUNVIDIA / CUDAModels & quantsWindowsLinux

Description

Part 2 of the #285 split. **One commit on 0.1.38** (part 1, #357, is in 0.1.38). The start times below were measured on 0.1.31; the checksums on 0.1.38 are in the comments.

0.1.31 registers the whole arena with CUDA before the first byte is read. With 4 KB pages, that registration is ~1.5 s of the start that nothing hides (IQ2_XS: 33 GiB). Every account without the Lock Pages privilege gets 4 KB pages.

What changes:
- **`PinnedArena(bytes, bounds, Deferred)` only reserves the arena.** `register_slices()` registers the per-layer slices in order on a thread and publishes how many are done.
- **Every reader writes layer L only once slices L and L+1 are registered.** This covers `load_experts_gguf` (buffered and unbuffered), `load_experts_ranges` and `load_experts_direct`.
  - Writing a page while `cudaHostRegister` runs on it corrupted the arena under WDDM, even on large pages.
  - A layer boundary inside a page puts that page in both registrations, so the next slice must be done too.
  - In effect this gives the page-aligned bounds you asked for, without moving the slice bounds, which an expert must not straddle.
- **A slice CUDA refuses ends the registration.** The rest is locked resident, as in the sliced fallback.
- **It is used only where the whole arena would be registered anyway:**
  - no cap (`STRATA_ARENA_PIN_GIB`, multi-GPU under WDDM);
  - no shared arena file;
  - Windows only: Linux keeps #253's whole-arena pin unchanged.
- **`STRATA_DEFERRED_REGISTER=0`** restores the old order.
- **`STRATA_VERIFY_ARENA=1`** prints the loaded arena's FNV-1a checksum (for tests).

**Checksums match.** The old order and the deferred one give the same checksum across:
- 4 KB pages and large pages;
- the GGUF read in place and experts.bin;
- buffered and unbuffered reads.

## Start times

Setup:
- IQ2_XS without experts.bin, RTX 5090, PCIe 5 drive, `generate` of 16 tokens.
- **Cold:** the IQ2_XS files were pushed out of the file cache by reading 142 GB of other files.
- **Restart:** the same start again right away.
- Each cell shows the time until the arena is loaded / until the 16-token answer is done. "Parts 1+2" is this branch.

**128 GB** (the file cache can keep the files), two to three runs each:

| | cold | restart |
|---|---|---|
| 4 KB pages (`STRATA_NO_LARGEPAGES=1`), 0.1.31 | 12.4-13.0 / 16.3-16.8 s | 7.8-8.5 / 10.9-11.7 s |
| 4 KB pages, parts 1+2 | **11.1-11.5 / 14.9-15.4 s** | **5.5-5.7 / 8.5-8.9 s** |
| large pages, 0.1.31 | 17.6-18.7 / 21.1-22.4 s | 4.6-5.1 / 7.0-7.6 s |
| large pages, parts 1+2 | 17.6-19.9 / 21.3-23.6 s | 4.1-4.5 / 6.7-7.1 s |

The cold spread with large pages comes from the drive: the cache flush reads 142 GB from the same NVMe. A/B runs of this branch with `STRATA_DEFERRED_REGISTER=0` (19.9 s) and with `STRATA_UNBUFFERED_LOAD=0` (17.6 / 18.8 s) fall inside that same spread.

**An emulated 64 GB PC** (a locked-pages RAM ballast leaves 58 GiB):
- 4 KB pages: two runs each.
- Large pages: four runs each.

| | cold | restart |
|---|---|---|
| 4 KB pages, 0.1.31 | 13.0-13.2 / 17.0-17.2 s | 12.9-14.1 / 16.8-18.0 s |
| 4 KB pages, parts 1+2 | **7.6-8.2 / 11.1-12.1 s** | **6.0-6.1 / 8.7 s** |
| large pages, 0.1.31 | 15.9-21.8 / 19.6-25.5 s | 12.2-22.6 / 15.7-26.2 s |
| large pages, parts 1+2 | 12.0-15.5 / 15.8-19.1 s | 6.1-11.9 / 8.7-14.7 s |

With large pages under that memory pressure, the allocation itself takes 6-12 s and varies from run to run. Windows has to gather contiguous 2 MB pages before the first byte is read, hence the wide ranges in both rows.

## Part 3 is dropped

I built and measured part 3 too: the arena opened on its own thread beside the dense weights, plus a cuBLAS handle made on a thread. It gave no net gain here:
- **The arena finished later:** 5.5 vs 4.5 s with large pages, 5.5-5.7 vs 4.9 s with 4 KB pages. The arena thread and cuBLAS init contend with the arena's readers and its registration.
- **The first decode step did not move:** it came at 7.1-7.7 s either way.

So it is not part of the split.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

Sur le site

Liens install, modèles, releases.