Issues / #1254

#1254 [Optimization] Support CPPC / configurable core placement for host loop and expert pool

open · @bceenaeiklmr · 0 コメント · GitHub で見る

BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

本文

### Feature / Optimization: Support CPPC preferred core placement for host loop and expert pool (avoid pinning host loop to slowest core)

#### Context
In `src/kernels/cpu/pool.cpp`, `CorePlacement::CorePlacement` initializes the placement for the host loop and the worker pool:

```cpp
CorePlacement::CorePlacement(bool spare_first) : cores_(physical_cores()) {
    for (size_t c = 0; c < cores_.size(); ++c)
        for (int p : cores_[c]) {
            if (p >= (int) core_of_.size()) core_of_.resize((size_t) p + 1, -1);
            core_of_[(size_t) p] = (int) c;
        }
    host_ = cores_.back().front();
    for (size_t c = spare_first && cores_.size() > 4 ? 1 : 0; c + 1 < cores_.size(); ++c)
        workers_.push_back(cores_[c].front());
}
```

#### Problem
On modern multi-core x86 processors with asymmetric core performance (AMD Zen 3/4/5 ACPI CPPC preferred cores, Intel hybrid P/E architectures), silicon binning varies across physical cores on the die.

Hardcoding `host_ = cores_.back().front()` blindly pins the latency-critical host thread to the highest-numbered physical core (e.g. Core 7 on an 8-core CPU). 

On an AMD Ryzen 7 5800X3D (8C/16T, dual RTX 3090 setup), querying the Windows kernel ACPI CPPC power capabilities (Kernel-Processor-Power Event ID 55) reveals:
* **Core 2 (Processors 4, 5)**: **158% max performance capability** (CPPC Preferred / Gold Star)
* **Core 3 (Processors 6, 7)**: **158% max performance capability** (CPPC Preferred / Gold Star)
* **Core 1 (Processors 2, 3)**: **154%**
* **Core 4 (Processors 8, 9)**: **150%**
* **Core 0 (Processors 0, 1)**: **145%** (intentionally spared for GPU PCIe interrupts and OS DPCs)
* **Core 5 (Processors 10, 11)**: **141%**
* **Core 6 (Processors 12, 13)**: **137%**
* **Core 7 (Processors 14, 15)**: **133% (the slowest, lowest-binned core on the entire chip)**

Because the host thread runs `session_loop`, spins on `cudaEventQuery`, schedules CUDA kernels across GPUs, conducts MTP verification, and handles token sampling, pinning it to the slowest silicon core caps its single-threaded boost frequency and introduces unnecessary scheduling latency.

#### Proposed Solution
1. **Configurable Environment Variables**:
Allow runtime overrides via environment variables (`STRATA_HOST_PROC` and `STRATA_WORKER_PROCS`) in `CorePlacement::CorePlacement`:
```cpp
const char* env_host = std::getenv("STRATA_HOST_PROC");
if (env_host != nullptr && *env_host != '\0') {
    int hp = std::atoi(env_host);
    if (hp >= 0 && hp < (int) core_of_.size()) host_ = hp;
}
const char* env_workers = std::getenv("STRATA_WORKER_PROCS");
if (env_workers != nullptr && *env_workers != '\0') {
    std::vector<int> w;
    std::stringstream ss(env_workers);
    std::string item;
    while (std::getline(ss, item, ',')) {
        int p = std::atoi(item.c_str());
        if (p >= 0 && p < (int) core_of_.size()) w.push_back(p);
    }
    if (!w.empty()) workers_ = w;
}
```

This allows setting in `strata-<model>.json`:
```json
"env": {
  "STRATA_HOST_PROC": "4",
  "STRATA_WORKER_PROCS": "6,2,8,10,12,14"
}
```

2. **Automated CPPC Detection (Optional Future Extension)**:
Alternatively, query OS CPPC information at startup:
- Windows: `GetSystemCpuSetInformation` / ACPI `_CPC` objects
- Linux: `/sys/devices/system/cpu/cpu*/acpi_cppc/highest_perf`
and sort cores by performance before allocating `host_` and `workers_`.

#### Verification & Benchmarks
Testing with Qwen3.8-Flash-Next on Dual RTX 3090s and Ryzen 7 5800X3D:
- By sparing Core 0 (0,1) for GPU interrupts, pinning the Host Thread to Core 2 (processor 4, 158% Gold Star), and assigning workers to Cores 3, 1, 4, 5, 6, 7 (`6,2,8,10,12,14`):
  - Short prompt prefill improved by **+6.0%** (58.7 tok/s -> 62.2 tok/s).
  - Host event loop and CUDA synchronization runs on the highest-clocked core instead of the silicon's worst core.

関連リンク

インストール・モデル・リリースへの站内リンク。