贡献 / #1598
#1598 `--aux-cpus`: +10% decode on a layer split by keeping helper threads off the pool's cores (Linux, off by default)
open · @noon-at-cgn · 0 评论 · 去 GitHub 看
BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
说明
## Title
Issue: Related: none found (searched 2026-10-07 and 2026-10-08). Nearby: #1166 moves the host thread to its SMT sibling, #501 / #626 pin the host, #1436 gives replicas disjoint core masks; none places the engine's other threads.
## Summary
**Why it matters:** the host thread is the one that answers every layer's doorbell during decode. Threads the engine starts later inherit its single CPU, and the CUDA driver's threads float over every CPU, the pool workers' included, so the host gets preempted in the middle of a window. Moving those threads to spare CPUs gave +10% decode on our box with no change to any result:
| 2x RTX 3080, layer split, 2 slots | `--aux-cpus` off | `auto` |
|---|---|---|
| solo decode (median of 12, two restarts) | 72.1 / 72.4 tok/s | 79.7 / 80.4 |
| two streams, aggregate (median of 10 rounds) | 80.8 / 82.8 tok/s | 90.8 / 90.7 |
| host thread involuntary context switches over the bench | 43.8k / 42.8k | 3.9k / 2.8k |
| prompt read (24k / 104k) | unchanged | unchanged |
`--aux-cpus off|auto|LIST` (or `STRATA_AUX_CPUS`): `auto` takes the host core's SMT siblings, then CPUs of physical cores with no worker, never one that shares a core with a worker. It moves threads only, never a result. Default off: the hooks do nothing.
## What changed
- `include/strata/platform/aux_cpus.hpp` (new, header only): spare-CPU derivation, list parser, `pin_current_thread()`, a sweep over `/proc/self/task` for the driver's threads, a way back to the old mask. Workers and host are never moved. Empty off Linux.
- `src/program/generate.cpp`: flag, variable, start-up plan, per-request switch; pins the adaptive tier's threads, stdin reader, watchdog.
- `expert_source.cpp`, `ple_reader.cpp`, `direct_file.cpp`, `prefill.cpp`: pin the look-ahead, PLE and file readers, prefill helpers, Stager copy threads (`STRATA_AUX_STAGER=0` leaves those). `session.cpp`, `pool.cpp`: mark host and workers.
- `serve/server.py`, `serve/test_server.py`: `strata_tune {"aux_cpus": 1|0}` is forwarded.
- `tests/platform/aux_cpus_test.cpp` + `CMakeLists.txt`; `docs/DETAILS.md`: section "Thread placement".
## Extra Notes
Measured on engine 0.1.40.3 (d5ea713) plus this change; the branch itself is cut from main fb58e0d.
Machine: 2x RTX 3080 20 GB (220 W), Xeon E5-2696 v4 (16 pool workers, host on CPU 2, helpers on CPU 24), 24 logical CPUs in the container, UD-Q4_K_XL, layer split 23, `--batch-mtp`, adaptive tier on. Same binary in every arm, only the flag differs.
- Restarts, auto / flag removed, ABAB: off/auto ratios solo 0.904 and 0.900, c2 0.890 and 0.912. Same config restarted twice: solo 1.009 / 1.005, c2 0.999 / 1.024.
- Live, no restart, per-request toggle (n=20 solo, 16 c2 rounds, rotated order): off/on 0.888 solo and 0.917 c2, against an A/A control of 0.985 and 0.990.
- doorbell -> flag A: mean and p99 are not worse without the flag; what appears is rare stalls above 100 us in 7-8 of 12 requests, against 0-1 with `auto`. That preemption causes the loss is our inference; the counters do not prove it.
- Warm-up discards are listed exactly in our notes (15 short requests + 4 decodes + 2 reads per restart arm).
Checked: `aux_cpus_test`, pool and direct-file CPU tests, `python -m unittest discover -s serve` (594 tests), CUDA build. Not built or run: HIP, SYCL, Windows, macOS. The no-spare-CPU path was not tested. `STRATA_ADAPT_JOB_CPU` still wins for the adaptive tier's job thread. Not covered: the Python launcher, the host thread.
本站相关内容
相关页面的快捷入口。