Pull requests / #1436
#1436 fix(serve): isolate CPU cores for independent Linux replicas
open · @agorevski · 0 comments · View on GitHub
BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPModels & quantsDocumentationLinux
Description
## Title <!-- If Applicable, reference the GitHub issue --> Issue: Not linked to an existing issue. Isolate CPU cores for independent Linux server replicas ## Summary <!-- Quick Summary of changes --> Independent one-GPU servers can select overlapping host and expert-worker cores. Reducing their worker counts alone does not prevent that contention. Give each selected replica a disjoint physical-core affinity mask, preserving allowed SMT siblings and the launcher's existing CPU restriction. **This PR adds the entire reference `setup_multigpu.sh` launcher.** It was a pre-existing, locally staged addition, absent from upstream base `d5ea7133`; this is not just an affinity patch to a launcher already shipped upstream. The PR is independent of engine optimizations and does not include local configuration files or installed-engine changes. ## What changed <!-- Specifics on files changed, and what changes were made there --> - `serve/cpu_affinity.py`: group allowed CPUs by Linux sysfs physical-package/core IDs, then partition available physical cores fairly. Keep allowed SMT siblings together without adding CPUs outside the mask. A single instance without reservations retains the full allowed mask. Reserve entire physical cores occupied by running unselected managed server PIDs. - `setup_multigpu.sh` (new executable): publish the full current root-relative reference launcher with its existing GPU selection, PID checks/restarts, API-key handling, 200 W power limit, health polling, and log streaming. Add per-server `taskset` masks, fail CPU allocation before stopping any servers when insufficient cores remain, and use `flock` to serialize allocation and spawning. Unselected managed servers remain running; children do not inherit the allocation lock. - `serve/test_cpu_affinity.py`: 10 tests cover fair 18-core partitioning, restricted masks, socket/core IDs, reserved siblings, single-instance behavior, insufficient cores/missing topology, inherited server masks, concurrent-launch rejection, and unselected managed servers. Launcher fixtures mock GPU, power, network, and log-tail operations; they use this worktree for scratch files and stop fixture servers before removing them. - `.gitignore`: ignore the root `.strata-multigpu.lock` file. - `docs/DETAILS.md`: add the independent Linux server CPU-contention paragraph, reference-launcher prerequisites, measured core-selection example, and limitations. ## Extra Notes <!-- Any extra notes, delete if there are none --> ### Scope and prerequisites The CPU allocator is generalized to Linux physical-package/core topology; the launcher is deliberately **not** a general fleet manager. It retains the local four-NVIDIA-GPU IQ3_S replica conventions: GPU IDs 0-3, default selection of all four, ports 8080-8083, and existing root-level `strata-iq3_s_gpu0.json` through `strata-iq3_s_gpu3.json` for whichever GPUs are selected. These model configuration files are not supplied by this PR. Users also need the configured models/engine, the root `.venv/bin/python` environment, Bash, Linux sysfs and process affinity support, `taskset`, `flock`, `ss`, `curl`, `nvidia-smi`, and the usual shell/log utilities. On each launcher startup, selected GPUs are set to **200 W**, using `sudo nvidia-smi` when not root. Hardware must support that limit, and permissions must allow it; `sudo` may prompt. No power-limit value was changed. Selected managed servers are stopped before the power-limit operation; a later power-limit failure prevents new server launches but does not restore stopped servers or already changed limits. The default binding is `0.0.0.0`. Before any server binds, the launcher requires a non-empty, newline-free API key, taking it from `--api-key`, `STRATA_API_KEY`, the key file, or generating one, then always passing `--api-key` to the server. The default key file is `$HOME/.config/strata/api-key`, overridable with `STRATA_API_KEY_FILE`. Use `--host 127.0.0.1` for local-only access. No actual keys, local JSON configurations, logs, PID files, `claude.sh`, or `launch.sh` are included. ### Behavior on the next launcher startup Existing servers are not modified by applying this PR. Affinity applies when the launcher next starts/restarts the selected instances. It excludes physical cores reserved by running unselected managed instances recorded in matching PID files; unrelated processes and unmanaged servers are not isolated by this mechanism. If insufficient free cores remain, allocation fails before stopping servers; restart the managed instances together or expand the launcher's allowed mask. The lock covers allocation and spawning, not the subsequent health-polling/log-streaming lifetime. On an 18-physical-core, 36-thread Xeon W-2295, four selected servers receive 5/5/4/4 physical cores. A previously observed real single GPU-3 IQ3_S engine under mask `14,15,16,17,32,33,34,35` selected host CPU 14 and worker CPUs 15-17; the SMT siblings 32-35 were also inherited. This observation is not a new fleet benchmark for this PR. A single selected server without reservations instead keeps the full allowed mask. The benefit is removal of overlapping physical-core affinity among managed replicas, not a claimed fleet-throughput improvement. Shared DRAM and memory-bandwidth contention remain. ### Validation Run from the independent `cpu-isolation` worktree: - `/home/algore/GIT/strata/.venv/bin/python -m unittest serve.test_cpu_affinity` — **passed**, 10 tests, no skips. - `bash -n setup_multigpu.sh` — **passed**. - `git diff --check` — **passed**. - `git diff --cached --check` — **passed**. - Executable-bit check and comparison with the current local launcher/helper — **passed**; launcher is committed as mode `100755`. No fixture scratch directories remained. **GPU fleet not tested.** The real launcher was not run; no GPU workload, real network/power operation, server restart, or power-limit change was performed. The original dirty checkout, source configuration files, and installed engines were left unchanged.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.