反馈 / #798
#798 Hybrid CPU detection on Linux counts only the Turbo Boost Max 3.0 "favored" cores as P-cores (Core Ultra 7 270K Plus: 2P/22E instead of 8P/16E) — affects setup's `--pool-workers` and `--pool-affinity auto|p-cores`
closed · @WUZICANGJIE · 1 评论 · 去 GitHub 看
BenchmarksSetup & installAMD / HIPWindowsLinux
说明
### Summary On Linux both `setup.py` (`cpu_cores()`, line 371) and the engine (`src/kernels/cpu/pool.cpp`, lines 221 and 253) classify a core as a P-core only when its `cpu_capacity` **equals** the maximum. On Intel CPUs with Turbo Boost Max 3.0, the kernel gives the two "favored" P-cores a slightly higher capacity than the other P-cores, so only those two are counted as P-cores and the remaining six P-cores are treated as E-cores. ### Environment - Intel Core Ultra 7 270K Plus (Arrow Lake, 8 P-cores + 16 E-cores, no SMT), Linux 7.2.8 (NixOS host, Strata run in an Ubuntu 24.04 container) - Strata `6f32ec0` (engine 0.1.39), HIP build, RX 9070 (gfx1201) ### Data ``` cpu_capacity: cpu0:1024 cpu1:1012 cpu2:1024 cpu3-7:1012 cpu8-23:768 /sys/devices/cpu_core/cpus: 0-7 /sys/devices/cpu_atom/cpus: 8-23 cpuinfo_max_freq: cpu0,2: 5.5 GHz cpu1,3-7: 5.4 GHz cpu8-23: 4.7 GHz ``` ### What happens - `setup.cpu_cores()` returns `(2, 22)`, so `hybrid_pool_workers()` recommends **12** workers instead of the **15** the #642 rule intends for 8P + 16E (`p - 1 + e // 2`). - The engine (0.1.38 and 0.1.39) logs `hybrid CPU detected (2 P-cores / 2 threads, 22 E-cores)`. With `--pool-affinity p-cores` or `--pool-affinity auto` and no explicit `--pool-workers`, it starts **1** expert-pool worker (`pool workers: 1`). - `--pool-affinity auto` with an explicit `--pool-workers 16` still lands on CPUs 1–16 (the misclassified P-cores 1, 3–7 sort first by CPU number), so it happens to work. ### Measured effect (decode, IQ2_XS, greedy 400 tokens on a cached 7K prompt, 4 runs after 2 warm-up runs, engine 0.1.38) | pool setting | workers | decode tok/s (mean) | |---|---|---| | default (`all`, one per core) | 23 | 51.7 | | `--pool-affinity auto` (default count) | 1 | 63.4 | | `--pool-affinity p-cores` (default count) | 1 | 62.4 | | `--pool-workers 12` (default affinity) | 12 | 66.6 | | `--pool-affinity auto --pool-workers 16` | 16 | 67.0–68.0 (two sessions) | On this machine the GPU is the bottleneck once the E-cores are out of the way (8, 12, 14 and 16 workers with the default affinity all gave 66.5–66.8 tok/s), so the wrong count costs little here. On a machine where the CPU half dominates, 1 worker (`p-cores`/`auto` defaults) or a too-small recommendation would likely cost more. ### Suggested fix - On Linux, prefer the hybrid PMU lists when they exist: `/sys/devices/cpu_core/cpus` (P) and `/sys/devices/cpu_atom/cpus` (E). - Otherwise, treat capacities close to the maximum as P-cores (for example `cap >= 0.9 * max_cap`) instead of `cap == max_cap`. Here that separates 1024/1012 (P) from 768 (E). - Windows (`EfficiencyClass`) is not affected. I have not checked whether other Turbo Boost Max 3.0 CPUs (e.g. Raptor Lake) report unequal P-core capacities in the same way. _This report was drafted with an AI assistant; all numbers above were measured on my machine._
本站相关内容
相关页面的快捷入口。