Pull requests / #9

#9 CPU pool: sleep between requests instead of spinning (fixes #4)

closed · merged 2026-09-26 · @Mirtraxxx · 0 commentaires · Sur GitHub

BenchmarksServer & APINVIDIA / CUDAWindowsLinux

Description

Fixes #4.

## What it does

A parked expert-pool worker waited for work with an endless `_mm_pause` spin on `epoch_`, so every worker's core stayed
at 100% while the engine sat idle between requests. On a Ryzen 7 5700X that's 7 cores, which the issue reports as
"CPU 50% even when doing nothing".

Now a parked worker spins exactly as before for up to 20 ms (`ExpertPool::kSpinBeforeSleep`), then blocks on a
condition variable:

- Every publish (`run`, `run_phase`, the destructor) goes through a new `publish()`. It bumps `epoch_` and notifies
  only if a worker is actually asleep, so the token path pays one uncontended load and never touches the mutex.
- Lost wakeups are ruled out by seq_cst on both sides. The worker does `sleepers_++` and then reads `epoch_` under the
  mutex. The host does `epoch_++` and then reads `sleepers_`. At least one of them sees the other's write. On x86 the
  `fetch_add` is a locked `xadd` either way.
- The clock is read once every 1024 pauses, so the spin itself is unchanged.
- Inside a request the gaps between batches are microseconds, so workers don't sleep mid-token. During a long GPU-only
  stretch of a prefill they can; waking costs microseconds against chunks of hundreds of milliseconds.

## Measured

RTX 3090, Ryzen 7 5700X (7 pool workers + host), 64 GB DDR4, Windows 11, Swift 1.5 IQ2_XS, `--serve`,
`--expert-cache auto --spec 4 --adapt-swaps 0`. Old and new engine, same flags, each run twice. CPU is process CPU
time over wall time, in cores.

| | old | new |
|---|---:|---:|
| Idle before any request (10 s) | 6.98 / 6.41 cores | **0.00 / 0.00** |
| Idle after requests (10 s) | 6.95 / 6.92 cores | **0.00 / 0.01** |
| Decode, 2.3K-token code prompt | 61.8 / 62.9 tok/s | 64.1 / 62.8 tok/s |
| Decode, short math prompt (384 tokens) | 52.5 / 52.0 tok/s | 50.3 / 54.3 tok/s |
| Decode, short writing prompt | 47.1 / 48.5 tok/s | 50.6 / 48.0 tok/s |
| Prefill, 2,298 tokens | 9.15 / 9.49 s | 8.65 / 8.59 s |

Decode speed is unchanged within run-to-run noise. The 2.3K-token prefill came out ~7% faster in both runs, likely
because the workers sleep through the GPU-only stretches and stop competing with the host for cores.

Greedy output: the code prompt is identical across all four runs. The other two diverge late in some runs, but the
old engine also diverges from itself between runs (math at token 53), so that's the existing run-to-run variation,
not this change. The pool's arithmetic isn't touched.

Only tested on Windows / one machine. The Linux path uses the same `std::condition_variable`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Sur le site

Liens install, modèles, releases.