Pull requests / #272

#272 feat(cpu): add hybrid architecture awareness and --pool-affinity for P/E-core CPUs

closed · @praveshkhatana · 0 comentários · No GitHub

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

Descrição

### Summary
On hybrid CPU architectures (e.g. Intel 12th–14th Gen / Core Ultra Alder Lake, Raptor Lake, Arrow Lake, and hybrid Linux systems), the CPU expert pool previously treated all physical cores as homogeneous and populated worker affinities linearly.

This created two issues on hybrid platforms:
1. **Worker Throttling:** When worker threads are assigned to slower Efficient cores (E-cores) alongside Performance cores (P-cores), barrier synchronization (`wait_done` / `wait_parked`) causes fast P-core workers to stall waiting for E-core workers, lowering aggregate CPU drain bandwidth.
2. **Resource Contention:** Core sizing could schedule workers onto SMT siblings of other workers when pool worker count exceeded physical P-core count.

### What this PR introduces
- **Hybrid Topology Detection (`detect_cpu_topology`)**:
  - **Windows**: Uses `GetLogicalProcessorInformationEx(RelationProcessorCore)` to inspect `Processor.EfficiencyClass` and `LTP_PC_SMT` flags to identify P-cores vs E-cores and logical thread pairings.
  - **Linux**: Reads `/sys/devices/system/cpu/cpu*/cpu_capacity` and topology sysfs files to identify core types.
- **Smart Worker Allocation**:
  - Prioritizes physical P-cores, reserving primary P-core 0 for the host thread and primary P-cores 1..N for worker threads.
  - SMT siblings and E-cores are only used as overflow or when requested.
- **Automatic Sizing**:
  - On hybrid CPUs, when `pool_workers` is unspecified (`<= 0`), it defaults to `(p_cores - 1)` so that all workers run unthrottled on dedicated physical P-cores.
- **CLI Configuration**:
  - Adds `--pool-affinity auto|p-cores|all` (default `auto`).
  - `--pool-affinity auto`: Prioritizes P-cores, defaults pool workers to P-core count minus host core.
  - `--pool-affinity p-cores`: Strictly restricts workers to P-cores and their SMT siblings.
  - `--pool-affinity all`: Preserves legacy sequential physical core allocation.
- **Diagnostics**:
  - Logs detected hybrid topology (`%d P-cores / %d threads, %d E-cores`) and effective worker count and affinity mode.

### Verification & Benchmarks
- Tested on Intel Core i9-14900K (8P + 16E cores, 32 threads) + NVIDIA RTX 4090 under Windows 11 with MSVC & CUDA 13.3.
- **Performance:** CPU pool drain throughput increased significantly (+30% to +45% faster CPU phase), improving token generation rate from ~46–54 tok/s to ~68 tok/s on short contexts and sustaining 52–56 tok/s on 100k token prompts.
- **Compatibility:** 100% backward compatible on non-hybrid CPUs (AMD Ryzen / Intel non-hybrid / Apple Silicon / cloud VMs) where `is_hybrid == false` falls back to identical legacy behavior.

No site

Links install, modelos, releases.