Pull requests / #626
#626 cpu: preserve full thread affinity on Windows and Linux
closed · @Yasei-no-otoko · 0 评论 · 在 GitHub 查看
BenchmarksAMD / HIPNVIDIA / CUDAWindowsLinux
描述
Strata's CPU pinning lost part of the processor identity on Windows and part of the caller's saved CPU mask on Linux. Windows workers discarded the processor-group number; Linux host restoration retained only CPU IDs 0-63. This change preserves the full placement on both platforms without changing automatic worker sizing or explicit `--pool-workers` values. On Windows, use `SetThreadGroupAffinity` for pool-owned workers, enumerate every core group mask and preserve group identity in fallback enumeration. On the tested 3990X, the previous 63-worker list collapsed to 32 distinct group-local masks because `SetThreadAffinityMask(..., 1 << (core & 63))` discarded the group. The caller-owned Windows host thread uses reversible CPU Set selection. Save its full explicit selection and restore it, or clear it when originally empty. This preserves process-default inheritance and Windows 11's implicit all-group eligibility without changing existing hard affinity. CPU Sets are soft placement constraints. See [Microsoft's CPU Set API](https://learn.microsoft.com/en-us/windows/win32/api/processthreadsapi/nf-processthreadsapi-setthreadselectedcpusets). On Linux, retain the complete native affinity mask in `ThreadAffinity`. Use `CPU_ALLOC` and the sized CPU-set macros, retry a query with a larger buffer on `EINVAL`, and share the dynamic-mask handling with topology enumeration and worker pinning. Restore the complete saved mask, including high-only and sparse selections, and report failed affinity operations using pthread's returned error number. See [Linux CPU-set APIs](https://man7.org/linux/man-pages/man3/CPU_SET.3.html). No Windows-style group IDs are needed on Linux. Validation: - Windows 11 Pro 10.0.26300 / Threadripper 3990X: rebuilt the Windows `pool_affinity_test` after the Linux addition; CTest passed. It covers worker pins, reversible host selections, process defaults, and continued execution in both processor groups. The earlier Windows change also completed a full Release HIP engine build; details are in #628. - Ubuntu 24.04 under WSL 2, GCC 13.3: built `strata_kernels_cpu` and `pool_affinity_test` with GPU/native-ggml support disabled. `-Wall -Wextra -Werror` syntax checks passed for the changed Linux implementation and test. CTest passed with 64 available CPUs, and a separate run with `taskset -c 0,2,15` passed. The high-CPU path correctly reports a skip on these WSL runs. - A diskless QEMU TCG guest using Linux `6.6.36.3-microsoft-standard-WSL2` exposed CPUs 0-127. The same statically linked regression test passed with `--require-high-cpu`, both with all 128 CPUs and with affinity restricted to CPUs 64-127. Tests include full-mask and sparse/nested restoration, a high-only saved mask, topology respecting the allowed CPU set, worker execution on CPU 127, invalid targets, and cleanup. - A separate negative-control binary deliberately truncated the saved Linux mask to one 64-bit word. Both QEMU scenarios failed as expected: the full mask restored only 64 of 128 CPUs, and the high-only scenario retained just the pinned CPU after restoration failed. The injected control is not part of the committed source. Linux reproduction on a normal Linux checkout: ```sh cmake -S . -B build-affinity -DSTRATA_BUILD_TESTS=ON \ -DSTRATA_NATIVE_EXPERTS=OFF -DSTRATA_ENABLE_CUDA=OFF -DSTRATA_ENABLE_HIP=OFF cmake --build build-affinity --target pool_affinity_test --parallel 4 ctest --test-dir build-affinity -R '^pool_affinity_test$' --output-on-failure -V # Requires at least two allowed CPUs, including an ID >= 64: ./build-affinity/pool_affinity_test --require-high-cpu ``` The Linux results establish affinity correctness in WSL / an emulated guest, not native Linux inference performance. CPU IDs above 1023 were not exercised. The AVX-512 pool self-test and `pool_stress` skip their workloads on this 3990X; a zero exit code from those skips is not a stress-workload pass. Benchmark data remains in the separate PR #628. Its speed figures belong to the recorded earlier Windows revisions and predate both the Windows host CPU Set review follow-up and this Linux addition; inference throughput was not remeasured for this revision. The machine's one NUMA node spans two processor groups, so this is an application affinity correction, not a Windows Containers or OS NUMA fix.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。