Issues / #798

#798 Hybrid CPU detection on Linux counts only the Turbo Boost Max 3.0 "favored" cores as P-cores (Core Ultra 7 270K Plus: 2P/22E instead of 8P/16E) — affects setup's `--pool-workers` and `--pool-affinity auto|p-cores`

closed · @WUZICANGJIE · 1 comentários · No GitHub

BenchmarksSetup & installAMD / HIPWindowsLinux

Descrição

### Summary

On Linux both `setup.py` (`cpu_cores()`, line 371) and the engine (`src/kernels/cpu/pool.cpp`, lines 221 and 253) classify a core as a P-core only when its `cpu_capacity` **equals** the maximum. On Intel CPUs with Turbo Boost Max 3.0, the kernel gives the two "favored" P-cores a slightly higher capacity than the other P-cores, so only those two are counted as P-cores and the remaining six P-cores are treated as E-cores.

### Environment

- Intel Core Ultra 7 270K Plus (Arrow Lake, 8 P-cores + 16 E-cores, no SMT), Linux 7.2.8 (NixOS host, Strata run in an Ubuntu 24.04 container)
- Strata `6f32ec0` (engine 0.1.39), HIP build, RX 9070 (gfx1201)

### Data

```
cpu_capacity:  cpu0:1024 cpu1:1012 cpu2:1024 cpu3-7:1012 cpu8-23:768
/sys/devices/cpu_core/cpus: 0-7        /sys/devices/cpu_atom/cpus: 8-23
cpuinfo_max_freq: cpu0,2: 5.5 GHz  cpu1,3-7: 5.4 GHz  cpu8-23: 4.7 GHz
```

### What happens

- `setup.cpu_cores()` returns `(2, 22)`, so `hybrid_pool_workers()` recommends **12** workers instead of the **15** the #642 rule intends for 8P + 16E (`p - 1 + e // 2`).
- The engine (0.1.38 and 0.1.39) logs `hybrid CPU detected (2 P-cores / 2 threads, 22 E-cores)`. With `--pool-affinity p-cores` or `--pool-affinity auto` and no explicit `--pool-workers`, it starts **1** expert-pool worker (`pool workers: 1`).
- `--pool-affinity auto` with an explicit `--pool-workers 16` still lands on CPUs 1–16 (the misclassified P-cores 1, 3–7 sort first by CPU number), so it happens to work.

### Measured effect (decode, IQ2_XS, greedy 400 tokens on a cached 7K prompt, 4 runs after 2 warm-up runs, engine 0.1.38)

| pool setting | workers | decode tok/s (mean) |
|---|---|---|
| default (`all`, one per core) | 23 | 51.7 |
| `--pool-affinity auto` (default count) | 1 | 63.4 |
| `--pool-affinity p-cores` (default count) | 1 | 62.4 |
| `--pool-workers 12` (default affinity) | 12 | 66.6 |
| `--pool-affinity auto --pool-workers 16` | 16 | 67.0–68.0 (two sessions) |

On this machine the GPU is the bottleneck once the E-cores are out of the way (8, 12, 14 and 16 workers with the default affinity all gave 66.5–66.8 tok/s), so the wrong count costs little here. On a machine where the CPU half dominates, 1 worker (`p-cores`/`auto` defaults) or a too-small recommendation would likely cost more.

### Suggested fix

- On Linux, prefer the hybrid PMU lists when they exist: `/sys/devices/cpu_core/cpus` (P) and `/sys/devices/cpu_atom/cpus` (E).
- Otherwise, treat capacities close to the maximum as P-cores (for example `cap >= 0.9 * max_cap`) instead of `cap == max_cap`. Here that separates 1024/1012 (P) from 768 (E).
- Windows (`EfficiencyClass`) is not affected.

I have not checked whether other Turbo Boost Max 3.0 CPUs (e.g. Raptor Lake) report unequal P-core capacities in the same way.

_This report was drafted with an AI assistant; all numbers above were measured on my machine._

No site

Links install, modelos, releases.