Issues / #1254
#1254 [Optimization] Support CPPC / configurable core placement for host loop and expert pool
open · @bceenaeiklmr · 0 コメント · GitHub で見る
BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux
本文
### Feature / Optimization: Support CPPC preferred core placement for host loop and expert pool (avoid pinning host loop to slowest core)
#### Context
In `src/kernels/cpu/pool.cpp`, `CorePlacement::CorePlacement` initializes the placement for the host loop and the worker pool:
```cpp
CorePlacement::CorePlacement(bool spare_first) : cores_(physical_cores()) {
for (size_t c = 0; c < cores_.size(); ++c)
for (int p : cores_[c]) {
if (p >= (int) core_of_.size()) core_of_.resize((size_t) p + 1, -1);
core_of_[(size_t) p] = (int) c;
}
host_ = cores_.back().front();
for (size_t c = spare_first && cores_.size() > 4 ? 1 : 0; c + 1 < cores_.size(); ++c)
workers_.push_back(cores_[c].front());
}
```
#### Problem
On modern multi-core x86 processors with asymmetric core performance (AMD Zen 3/4/5 ACPI CPPC preferred cores, Intel hybrid P/E architectures), silicon binning varies across physical cores on the die.
Hardcoding `host_ = cores_.back().front()` blindly pins the latency-critical host thread to the highest-numbered physical core (e.g. Core 7 on an 8-core CPU).
On an AMD Ryzen 7 5800X3D (8C/16T, dual RTX 3090 setup), querying the Windows kernel ACPI CPPC power capabilities (Kernel-Processor-Power Event ID 55) reveals:
* **Core 2 (Processors 4, 5)**: **158% max performance capability** (CPPC Preferred / Gold Star)
* **Core 3 (Processors 6, 7)**: **158% max performance capability** (CPPC Preferred / Gold Star)
* **Core 1 (Processors 2, 3)**: **154%**
* **Core 4 (Processors 8, 9)**: **150%**
* **Core 0 (Processors 0, 1)**: **145%** (intentionally spared for GPU PCIe interrupts and OS DPCs)
* **Core 5 (Processors 10, 11)**: **141%**
* **Core 6 (Processors 12, 13)**: **137%**
* **Core 7 (Processors 14, 15)**: **133% (the slowest, lowest-binned core on the entire chip)**
Because the host thread runs `session_loop`, spins on `cudaEventQuery`, schedules CUDA kernels across GPUs, conducts MTP verification, and handles token sampling, pinning it to the slowest silicon core caps its single-threaded boost frequency and introduces unnecessary scheduling latency.
#### Proposed Solution
1. **Configurable Environment Variables**:
Allow runtime overrides via environment variables (`STRATA_HOST_PROC` and `STRATA_WORKER_PROCS`) in `CorePlacement::CorePlacement`:
```cpp
const char* env_host = std::getenv("STRATA_HOST_PROC");
if (env_host != nullptr && *env_host != '\0') {
int hp = std::atoi(env_host);
if (hp >= 0 && hp < (int) core_of_.size()) host_ = hp;
}
const char* env_workers = std::getenv("STRATA_WORKER_PROCS");
if (env_workers != nullptr && *env_workers != '\0') {
std::vector<int> w;
std::stringstream ss(env_workers);
std::string item;
while (std::getline(ss, item, ',')) {
int p = std::atoi(item.c_str());
if (p >= 0 && p < (int) core_of_.size()) w.push_back(p);
}
if (!w.empty()) workers_ = w;
}
```
This allows setting in `strata-<model>.json`:
```json
"env": {
"STRATA_HOST_PROC": "4",
"STRATA_WORKER_PROCS": "6,2,8,10,12,14"
}
```
2. **Automated CPPC Detection (Optional Future Extension)**:
Alternatively, query OS CPPC information at startup:
- Windows: `GetSystemCpuSetInformation` / ACPI `_CPC` objects
- Linux: `/sys/devices/system/cpu/cpu*/acpi_cppc/highest_perf`
and sort cores by performance before allocating `host_` and `workers_`.
#### Verification & Benchmarks
Testing with Qwen3.8-Flash-Next on Dual RTX 3090s and Ryzen 7 5800X3D:
- By sparing Core 0 (0,1) for GPU interrupts, pinning the Host Thread to Core 2 (processor 4, 158% Gold Star), and assigning workers to Cores 3, 1, 4, 5, 6, 7 (`6,2,8,10,12,14`):
- Short prompt prefill improved by **+6.0%** (58.7 tok/s -> 62.2 tok/s).
- Host event loop and CUDA synchronization runs on the highest-clocked core instead of the silicon's worst core.
関連リンク
インストール・モデル・リリースへの站内リンク。