Issues / #1056

#1056 Prefill: Stager threads yield-spin while waiting. On Linux with --mmap-experts (DGX Spark / GB10), prompts are read ~10x slower; elsewhere ~20 cores are burned during prefill

closed · @uncle-daddy-jp · 1 commentaires · Sur GitHub

BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsLinux

Description

**Summary**

On a DGX Spark (GB10), using the `arm-gb10-port` branch from PR #409 (44df8fd), reading a prompt with `--mmap-experts` ran at ~90 tok/s instead of ~1,200-1,400 tok/s. The cause is the prefill `Stager`: its threads wait by spinning (`std::this_thread::yield()` on `issued`, and `cudaEventSynchronize` on events that spin under the default `cudaDeviceScheduleAuto`). With a transient expert source there are 32 of them. On Linux they inherit the host thread's single-core pin (`SessionLoopScratch::init` -> `pin_current_thread(physical_cores(false)[0])`), so all 32 spin on the host's core and starve it. If they are unpinned, prompts are fast again, but the 32 spinners then take every core (process CPU 1,700-2,000% on a 20-core box) for the whole prompt.

Making the two waits sleep fixes both the slowdown and the CPU use, and no affinity change is needed. The patch is 5 lines of code plus a 2-line comment; PR follows.

**Environment**

- NVIDIA DGX Spark (GB10, aarch64, 10x Cortex-X925 + 10x Cortex-A725, 128 GB unified memory), driver 580.142, CUDA 13.0, kernel 6.17.0-1014-nvidia, GCC 13.3, Ubuntu 24.04
- Strata `arm-gb10-port` 44df8fd (PR #409). `src/prefill/prefill.cpp` is unchanged between 99f3dbd and that branch; the same code is on main (1735d64 as of today), and the patch applies cleanly there (`git apply --check`).
- Model: Qwen3.8-Flash-Next IQ3_S (GSQ-RCO), `--mmap-experts --expert-cache auto --prefill auto --spec 4 --mtp ... --kv int8 --vision`, all 24,576 experts in the GPU cache.

**Repro**

Linux plus an expert source whose blobs are `transient()` (a GGUF read in place, `--mmap-experts`), so that `Prefill::init` starts 32 stager threads (`files ? 32 : ...`). Send an 8K or 30K-token prompt.

**Evidence (before)**

- `STRATA_PREFILL_TIMING=1`: 68-69% of prefill time was "host grouping" (~0.9 s per MoE layer). It drops to 0.5-1.4% with the fix.
- `/proc/<pid>/task/*/status`: 35-36 threads had `Cpus_allowed_list: 0`. On GB10, CPU 0 is an A725 (cpu_capacity 718), because `physical_cores(false)` with the default `PoolAffinity::All` lists CPUs in number order.
- `perf record -g` during prefill: ~85% of samples in `Stager::work` -> `sched_yield` / `__schedule`.
- `taskset -p -c 0-19` on every task of the running engine: prompt reading became 10x faster at once (IQ3_S 8K 91 -> 1,199 tok/s, 30K 88 -> 1,357 tok/s).
- After unpinning only the stagers: still fast, but `perf` showed the time going to `__schedule`/`do_sched_yield` (~45%) plus a libcuda spin in `cudaEventSynchronize` (~15%), and process CPU was 1,640-1,880% (peaks 2,000%) for the whole prompt.

**Results** (IQ3_S, prompts of 7,031 / 26,383 tokens, server-reported tok/s; CPU = process CPU sampled every 1 s, 100 = one core)

| build | 8K read | 30K read | CPU while reading | CPU idle | decode |
|---|---:|---:|---:|---:|---:|
| 44df8fd as is | 91 | 88 | ~100 (all on core 0) | 0 | 42-48 |
| stagers unpinned (affinity reset in `Stager::work`) | 1,068-1,183 | 1,344-1,397 | 1,640-1,880 | 0 | 40-42 |
| ... + `STRATA_STAGER_THREADS=8` | 1,237 | 1,397 | 846-900 | 0 | 41-42 |
| ... + `STRATA_STAGER_THREADS=4` | 1,234 | 1,397 | 475-500 | 0 | 42 |
| ... + ring wait as `atomic::wait` | 1,227 | 1,387 | 596-648 | 0 | 41-43 |
| ... + blocking-sync `dma_done` events | 1,185-1,230 | 1,342-1,399 | 112-118 | 0 | 40-43 |
| **the patch alone (no affinity change)** | **1,222-1,229** | **1,388-1,390** | **102** | **0** | **41-43** |

With the patch, the remaining ~1 core during prefill is the host thread and the issuer (`Stager::wait`, `wait_issued`), which still spin by design and share the host's pinned core. Decode is unchanged. The 32 copy threads can stay on core 0 without slowing anything down: the copies are a small part of the time (`__memcpy_sve` ~1-2% of samples) and the waiting threads now sleep.

**Patch**

`atomic::wait`/`notify_all` for the ring wait, `cudaEventBlockingSync` for `dma_done`, and one define in the HIP shim (not tested on HIP). Opened as a PR.

**Notes**

- The ring wait wakes all waiters on each `issued_one` (`notify_all`). With 32 threads that is cheap in practice (process CPU ~1,800% -> ~100%). A per-slot wait would avoid it if it ever matters.
- Separately from this patch: on Linux, any helper thread started by the pinned host thread (the stagers, the prefill issuer) inherits the host's single core. Harmless once they sleep, but worth knowing. On hybrid ARM (GB10), `physical_cores(false)[0]` picks a little core (CPU 0) for the host.
- Workaround without a rebuild: `STRATA_STAGER_THREADS=4` plus `taskset` (or any unpinning). With 4 threads the slowdown is gone on GB10, but 4 cores still spin.
- Measured with the help of Claude (Anthropic); every number above is from runs on the machine described.

Sur le site

Liens install, modèles, releases.