Pull requests / #1166
#1166 cpu pool: --host-core sibling, the host off the interrupts' logical processor (hybrid CPUs too)
open · @Hardin22 · 0 Kommentare · Auf GitHub
BenchmarksNVIDIA / CUDAModels & quantsWindowsLinux
Beschreibung
`--host-core last` exists because Windows sends a GPU's interrupts to the first logical processor, where the host thread spins on the GPU's flags. It leaves hybrid CPUs alone, though, so on a 12th-14th gen Intel or a Core Ultra the host stays on logical processor 0 together with the interrupts. This adds `--host-core sibling` (or `STRATA_HOST_CORE=sibling`). The host stays on the first core but moves to its other hardware thread: - the interrupts keep logical processor 0, and nothing else is pinned there; - the workers keep every other core, since they take one logical processor per core. With `--pool-affinity auto` / `p-cores`, which also use the P-cores' siblings, the host's sibling is taken out of their list; - it works on hybrid CPUs, and it is `first` when the core has no SMT sibling; - Linux gets the same layout (the sibling from sysfs), though there the interrupts are usually spread anyway. It is opt-in, so nothing changes by default. The startup log names the placement, e.g. `the host thread on 1 (draining too) (--host-core sibling)`. ## Measured RTX 4060 Ti (stage 0, layers 0-19, on PCIe 4.0 x4) + RTX 5080 (stage 1), i9-14900KF (8P+16E, hybrid), Windows 11. Swift 1.5 IQ3_XXS at 160K (q4_0 KV), resident RAM mode with the whole complement in RAM, stock draft layer, `--trim-stage-weights`. Greedy, five prompts × two rounds, a fresh server per run, the two arms interleaved, decode tok/s: | | `first` (today) | `sibling` | |---|---|---| | 0.1.40.1, `--pipeline-windows 2` | 98.0, 96.0, 98.2 | 101.4, 99.6, 98.2 | | 0.1.40.1 + #1122, `--pipeline-windows 2 --adapt-async 1` | 102.1, 112.5, 98.4 | 112.1, 112.2, 111.4 | The first row is about +2.4%. The second is +7% on average (104.3 → 111.9), and the spread goes from 14 tok/s to 0.8. That is where it matters most: there the host also queues the async tier's copies, and with the host on logical processor 0 a whole run would sometimes come out 10-15% slower, while the GPU clocks, P-states, acceptance and cache hits were the same. In the decode timing log, the slow runs had the same per-layer host time but more "GPU-reach wait". With `sibling` those drops went away: an earlier set of four pairs on the same stack, with some opt-in pipeline switches on, gave 105.5 ± 6 with the host on logical processor 0 and 114.2 ± 0.4 on its sibling. The build for the second row also carried other branches' opt-in switches, all off, and a Windows power-throttling opt-out in both arms, which made no difference on its own. ## Checks - `pool_affinity_test` (Windows) gains a layout check for every affinity mode (`all`, `auto`, `p-cores`): with `sibling`, the host is on the first host core's sibling, no worker is on that core, and without SMT the layout is the same as `first`. It passes on the 14900KF. The Linux test is unchanged, and the Linux path (`sibling_of` from sysfs) was neither built nor run here. - Builds with MSVC + CUDA 13.3 (sm_89 / sm_120). - Only threads move, so results are the same. If it holds on other hybrid boxes, it may be worth making it the default for hybrid CPUs on Windows. I left that to you.
Mehr auf der Site
Links zu Install, Modellen, Releases.