Pull requests / #611

#611 cpu: group ARM capacities and honor the pool's reserved host core

closed · @Suoriks · 0 comments · View on GitHub

NVIDIA / CUDAWindowsLinux

Description

On DGX Spark, Linux reports X925 CPU capacities of 997, 1017 and 1024, and A725 capacities of 718 and 731. Selecting only CPUs whose capacity equals the maximum counts one performance core instead of ten. The session also reselects its host CPU using the default `All` policy, even when the expert pool reserved a different CPU under `auto` or `p-cores`.

This change groups ARM capacities into two classes when their spread is at least 10%, using the midpoint as the boundary. Smaller variations or incomplete capacity data keep physical-core order. The existing x86 capacity classification and Windows topology discovery are retained. Linux layout selection is separated from sysfs reads so recorded topologies can be tested without matching hardware.

`ExpertPool::host_core()` exposes the CPU already reserved by the pool. `generate` passes it to `SessionLoopScratch::init`, which pins the host there and restores affinity at teardown. Existing two-argument callers retain the default host selection.

Validation:

- Windows x86-64, MSVC, C++17, `/O2 /DNDEBUG /W4 /WX`: `cpu_topology_test` passed. Fixtures cover recorded GB10 capacities, all three policies, restricted allowed CPUs, homogeneous capacity variation, missing capacity, SMT ordering, empty sets and the existing x86 classification.
- DGX Spark / GB10, Linux aarch64, GCC 13, CUDA 13, `sm_121`: the engine built in an isolated checkout with the previously validated local ARM port applied. `ctest -R '^(cpu_topology_test|session_host_affinity_test)$' -V` passed both tests. The session test used actual CUDA staging/event allocation and OS affinity queries: `all` reserved CPU 0; `auto` and `p-cores` reserved CPU 5 (X925). Each run restored the original affinity; the legacy init call also passed.

The ARM build support is covered by #409; this PR contains the topology and host-placement fixes. The two-class ARM heuristic has hardware coverage on GB10 only. Full x86 engine execution was not tested. No throughput improvement is claimed.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.