Issues / #921
#921 CPU expert-pool 20 ms spin can hurt GPU-heavy decode; 100 µs gives ~15% higher TG and ~4× lower CPU usage on dual RTX 4090
open · @hipotures · 4 Kommentare · Auf GitHub
BenchmarksMulti-GPUNVIDIA / CUDAModels & quants
Beschreibung
I found a potentially large performance issue with the default CPU expert-pool spin-before-sleep policy
on a GPU-heavy / mostly-resident configuration.
This is not a correctness issue. The existing `STRATA_POOL_SPIN_US` test knob is enough to reproduce
the effect, so I am filing this without a patch — the right upstream fix may be a different default,
adaptive parking, or another policy.
## System / configuration
- 2× RTX 4090 24 GiB
- Ryzen 9 7950X3D
- VM with 16 vCPUs
- Qwen3.8-Flash-Next IQ3_S
- current Strata v0.1.39-era source/binary
- layer split K=25
- `pcie-frac=0.28`
- 15 CPU expert-pool workers
- MTP `spec=4`, `spec-min-p=0.5`
- suffix lookup off
- prompt reuse off
- serial requests
- 4096 output tokens
The comparison below uses the **same executable**. The only relevant change is:
```bash
STRATA_POOL_SPIN_US=100
```
instead of the default 20 ms spin-before-sleep.
Each point has 3 valid measured requests after the same warmup.
## Result
| Total context | Default TG median | 100 us TG median | Change |
|---:|---:|---:|---:|
| 32K | 155.9 tok/s | **178.5 tok/s** | **+14.5%** |
| 128K | 133.2 tok/s | **154.0 tok/s** | **+15.6%** |
The individual 100 us TG runs were:
- 32K: `178.2 / 178.6 / 178.5`
- 128K: `144.1 / 154.0 / 165.4`
Median request wall time fell by approximately 11% at both context sizes.
## CPU utilization
The more surprising result is CPU usage during decode:
| Total context | Default VM CPU | 100 us VM CPU |
|---:|---:|---:|
| 32K | ~99.2% | **~28.0%** |
| 128K | ~97.9% | **~25.8%** |
GPU utilization remained roughly unchanged.
The generated output, MTP behavior, expert-cache capacities and normal routing counters were identical in the paired comparison, so the speed difference is not explained by a different greedy/MTP trajectory or by giving the candidate more cache.
## What seems to be happening
The current pool has:
```cpp
static constexpr std::chrono::milliseconds kSpinBeforeSleep{20};
```
and `STRATA_POOL_SPIN_US` overrides it.
On this configuration most expert work is already handled by the GPUs. Many layers therefore have little or no CPU expert work, but the worker pool can continue spinning for up to 20 ms after becoming idle.
With 15 workers this keeps almost the entire VM CPU busy even when those cores are not doing useful expert computation.
Reducing the spin period to 100 us lets the workers park much sooner and substantially reduces CPU contention.
This appears to improve GPU-side inference throughput as well, rather than merely reducing power/CPU usage.
## Important trade-off
100 us is not free.
In separate selected-layer diagnostics, CPU-positive completion-wait medians increased from roughly:
- **4–6 us** with the default policy
to roughly:
- **59–72 us** with the 100 us setting.
All-local wait floors remained around 4–5 us.
So I would not suggest simply changing the global default from 20 ms to 100 us without more workload coverage.
This probably depends on how often the CPU expert pool is actually needed.
## Possible upstream direction
An adaptive policy may be preferable to one fixed spin duration.
For example:
- park quickly after mostly/all-local GPU layers,
- retain a longer spin window when CPU expert work has recently occurred or jobs are pending,
- possibly adapt from recent CPU-positive frequency / completion latency.
The important observation is that VM-wide ~100% CPU in this configuration was mostly not useful expert computation: reducing the idle-spin window both lowered CPU use dramatically and increased decode throughput.
## Reproducibility / caveat
The clean benchmark used one long repository-maintenance workload at 32K and 128K total context limits.
The result therefore does **not** establish a universal +15% improvement yet, especially for workloads with substantially more CPU misses.
I am currently validating the 100 us setting on independent code/math/structured workloads, including more CPU-positive cases.
The current evidence is strong enough that I thought the fixed 20 ms default was worth reporting now.
Mehr auf der Site
Links zu Install, Modellen, Releases.