Pull requests / #1436

#1436 fix(serve): isolate CPU cores for independent Linux replicas

open · @agorevski · 0 コメント · GitHub で見る

BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPModels & quantsDocumentationLinux

本文

## Title
<!-- If Applicable, reference the GitHub issue -->
Issue: Not linked to an existing issue.

Isolate CPU cores for independent Linux server replicas

## Summary
<!-- Quick Summary of changes -->
Independent one-GPU servers can select overlapping host and expert-worker cores. Reducing their worker counts alone does not prevent that contention. Give each selected replica a disjoint physical-core affinity mask, preserving allowed SMT siblings and the launcher's existing CPU restriction.

**This PR adds the entire reference `setup_multigpu.sh` launcher.** It was a pre-existing, locally staged addition, absent from upstream base `d5ea7133`; this is not just an affinity patch to a launcher already shipped upstream. The PR is independent of engine optimizations and does not include local configuration files or installed-engine changes.

## What changed
<!-- Specifics on files changed, and what changes were made there -->
- `serve/cpu_affinity.py`: group allowed CPUs by Linux sysfs physical-package/core IDs, then partition available physical cores fairly. Keep allowed SMT siblings together without adding CPUs outside the mask. A single instance without reservations retains the full allowed mask. Reserve entire physical cores occupied by running unselected managed server PIDs.
- `setup_multigpu.sh` (new executable): publish the full current root-relative reference launcher with its existing GPU selection, PID checks/restarts, API-key handling, 200 W power limit, health polling, and log streaming. Add per-server `taskset` masks, fail CPU allocation before stopping any servers when insufficient cores remain, and use `flock` to serialize allocation and spawning. Unselected managed servers remain running; children do not inherit the allocation lock.
- `serve/test_cpu_affinity.py`: 10 tests cover fair 18-core partitioning, restricted masks, socket/core IDs, reserved siblings, single-instance behavior, insufficient cores/missing topology, inherited server masks, concurrent-launch rejection, and unselected managed servers. Launcher fixtures mock GPU, power, network, and log-tail operations; they use this worktree for scratch files and stop fixture servers before removing them.
- `.gitignore`: ignore the root `.strata-multigpu.lock` file.
- `docs/DETAILS.md`: add the independent Linux server CPU-contention paragraph, reference-launcher prerequisites, measured core-selection example, and limitations.

## Extra Notes
<!-- Any extra notes, delete if there are none -->
### Scope and prerequisites
The CPU allocator is generalized to Linux physical-package/core topology; the launcher is deliberately **not** a general fleet manager. It retains the local four-NVIDIA-GPU IQ3_S replica conventions: GPU IDs 0-3, default selection of all four, ports 8080-8083, and existing root-level `strata-iq3_s_gpu0.json` through `strata-iq3_s_gpu3.json` for whichever GPUs are selected. These model configuration files are not supplied by this PR. Users also need the configured models/engine, the root `.venv/bin/python` environment, Bash, Linux sysfs and process affinity support, `taskset`, `flock`, `ss`, `curl`, `nvidia-smi`, and the usual shell/log utilities.

On each launcher startup, selected GPUs are set to **200 W**, using `sudo nvidia-smi` when not root. Hardware must support that limit, and permissions must allow it; `sudo` may prompt. No power-limit value was changed. Selected managed servers are stopped before the power-limit operation; a later power-limit failure prevents new server launches but does not restore stopped servers or already changed limits.

The default binding is `0.0.0.0`. Before any server binds, the launcher requires a non-empty, newline-free API key, taking it from `--api-key`, `STRATA_API_KEY`, the key file, or generating one, then always passing `--api-key` to the server. The default key file is `$HOME/.config/strata/api-key`, overridable with `STRATA_API_KEY_FILE`. Use `--host 127.0.0.1` for local-only access. No actual keys, local JSON configurations, logs, PID files, `claude.sh`, or `launch.sh` are included.

### Behavior on the next launcher startup
Existing servers are not modified by applying this PR. Affinity applies when the launcher next starts/restarts the selected instances. It excludes physical cores reserved by running unselected managed instances recorded in matching PID files; unrelated processes and unmanaged servers are not isolated by this mechanism. If insufficient free cores remain, allocation fails before stopping servers; restart the managed instances together or expand the launcher's allowed mask. The lock covers allocation and spawning, not the subsequent health-polling/log-streaming lifetime.

On an 18-physical-core, 36-thread Xeon W-2295, four selected servers receive 5/5/4/4 physical cores. A previously observed real single GPU-3 IQ3_S engine under mask `14,15,16,17,32,33,34,35` selected host CPU 14 and worker CPUs 15-17; the SMT siblings 32-35 were also inherited. This observation is not a new fleet benchmark for this PR. A single selected server without reservations instead keeps the full allowed mask.

The benefit is removal of overlapping physical-core affinity among managed replicas, not a claimed fleet-throughput improvement. Shared DRAM and memory-bandwidth contention remain.

### Validation
Run from the independent `cpu-isolation` worktree:
- `/home/algore/GIT/strata/.venv/bin/python -m unittest serve.test_cpu_affinity` — **passed**, 10 tests, no skips.
- `bash -n setup_multigpu.sh` — **passed**.
- `git diff --check` — **passed**.
- `git diff --cached --check` — **passed**.
- Executable-bit check and comparison with the current local launcher/helper — **passed**; launcher is committed as mode `100755`. No fixture scratch directories remained.

**GPU fleet not tested.** The real launcher was not run; no GPU workload, real network/power operation, server restart, or power-limit change was performed. The original dirty checkout, source configuration files, and installed engines were left unchanged.

関連リンク

インストール・モデル・リリースへの站内リンク。