Pull requests / #9
#9 CPU pool: sleep between requests instead of spinning (fixes #4)
closed · merged 2026-09-26 · @Mirtraxxx · 0 comments · View on GitHub
BenchmarksServer & APINVIDIA / CUDAWindowsLinux
Description
Fixes #4. ## What it does A parked expert-pool worker waited for work with an endless `_mm_pause` spin on `epoch_`, so every worker's core stayed at 100% while the engine sat idle between requests. On a Ryzen 7 5700X that's 7 cores, which the issue reports as "CPU 50% even when doing nothing". Now a parked worker spins exactly as before for up to 20 ms (`ExpertPool::kSpinBeforeSleep`), then blocks on a condition variable: - Every publish (`run`, `run_phase`, the destructor) goes through a new `publish()`. It bumps `epoch_` and notifies only if a worker is actually asleep, so the token path pays one uncontended load and never touches the mutex. - Lost wakeups are ruled out by seq_cst on both sides. The worker does `sleepers_++` and then reads `epoch_` under the mutex. The host does `epoch_++` and then reads `sleepers_`. At least one of them sees the other's write. On x86 the `fetch_add` is a locked `xadd` either way. - The clock is read once every 1024 pauses, so the spin itself is unchanged. - Inside a request the gaps between batches are microseconds, so workers don't sleep mid-token. During a long GPU-only stretch of a prefill they can; waking costs microseconds against chunks of hundreds of milliseconds. ## Measured RTX 3090, Ryzen 7 5700X (7 pool workers + host), 64 GB DDR4, Windows 11, Swift 1.5 IQ2_XS, `--serve`, `--expert-cache auto --spec 4 --adapt-swaps 0`. Old and new engine, same flags, each run twice. CPU is process CPU time over wall time, in cores. | | old | new | |---|---:|---:| | Idle before any request (10 s) | 6.98 / 6.41 cores | **0.00 / 0.00** | | Idle after requests (10 s) | 6.95 / 6.92 cores | **0.00 / 0.01** | | Decode, 2.3K-token code prompt | 61.8 / 62.9 tok/s | 64.1 / 62.8 tok/s | | Decode, short math prompt (384 tokens) | 52.5 / 52.0 tok/s | 50.3 / 54.3 tok/s | | Decode, short writing prompt | 47.1 / 48.5 tok/s | 50.6 / 48.0 tok/s | | Prefill, 2,298 tokens | 9.15 / 9.49 s | 8.65 / 8.59 s | Decode speed is unchanged within run-to-run noise. The 2.3K-token prefill came out ~7% faster in both runs, likely because the workers sleep through the GPU-only stretches and stop competing with the host for cores. Greedy output: the code prompt is identical across all four runs. The other two diverge late in some runs, but the old engine also diverges from itself between runs (math at token 53), so that's the existing run-to-run variation, not this change. The pool's arithmetic isn't touched. Only tested on Windows / one machine. The Linux path uses the same `std::condition_variable`. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.