Pull requests / #650
#650 pinned: back the Linux expert arena with transparent huge pages
closed · @anon761 · 0 Kommentare · Auf GitHub
BenchmarksServer & APIMulti-GPUNVIDIA / CUDAModels & quantsWindowsLinux
Beschreibung
## What Without a hugetlb pool (`vm.nr_hugepages=0`, the default on most distros) the Linux expert arena falls back to a plain anonymous mapping with 4 KB pages: ``` strata generate: expert arena: cudaHostRegister PORTABLE ok; MAP_HUGETLB unavailable (needed 36727 2 MiB pages, vm.nr_hugepages=0); using 4 KB pages ``` The CPU expert pool streams whole experts (3-4 MB each for the 4-5 bit GGUFs) out of that arena, so every expert walks ~750-1000 pages through the TLB, and pinning ~72 GiB with `cudaHostRegister` has to lock ~18.8 M pages. This maps the fallback 2 MiB-aligned and `madvise(MADV_HUGEPAGE)`s it, so THP in `madvise` mode (the common distro default) backs it. `STRATA_NO_LARGEPAGES=1` keeps 4 KB pages, as on the hugetlb path. The log note says which one the arena got: ``` strata generate: expert arena: cudaHostRegister PORTABLE ok; MAP_HUGETLB unavailable (needed 36727 2 MiB pages, vm.nr_hugepages=0); transparent huge pages requested (MADV_HUGEPAGE) ``` Linux only; the Windows path and the hugetlb path are unchanged. ## Measured 2x RTX 3090 (layer split auto), EPYC 7413 (24C, AVX2), Unsloth UD-Q4_K_XL (71.7 GiB of experts), 262K context, `--kv int8 --pcie-frac 0 --ple-io ram`, engine at current `main`, same binary flags otherwise. Before = `main`, after = this branch. | | before | after | | --- | ---: | ---: | | engine start until `/v1/models` says loaded (several starts each) | 65 s | 40-45 s | | expert load (`loaded 71.73 GiB at ...`) | 3.56 GiB/s | 3.59 GiB/s | | decode, default expert cache (2K / 8K prompt) | 92.8 / 93.1 tok/s | 93.2 / 88.2 tok/s | | decode, `--expert-cache 1200` (CPU computes most misses), 4 runs | 69.5 tok/s | 71.1 tok/s | | CPU pool per expert entry (`STRATA_DECODE_TIMING`, same run) | 1.31 ms | 1.25 ms | The read itself is unchanged, so the shorter start is most likely the driver registering 2 MiB pages instead of 4 KB ones. With the default cache the CPU pool has little to do on this box (~1 ms per window), so decode is within run-to-run noise there; where the CPU computes most misses the pool's time per expert drops ~4-5%.
Mehr auf der Site
Links zu Install, Modellen, Releases.