Pull requests / #1057

#1057 prefill: let Stager threads sleep instead of yield-spinning (Linux + --mmap-experts: prompts ~10x faster on DGX Spark, ~1 core instead of ~20)

closed · @uncle-daddy-jp · 0 comentários · No GitHub

BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsLinux

Descrição

Fixes #1056.

Fixes the prefill slowdown and CPU burn described in the linked issue: the `Stager` threads waited by spinning (a `yield` loop on `issued`, and `cudaEventSynchronize` on events created without `cudaEventBlockingSync`, which spins under the default scheduler). With `--mmap-experts` there are 32 of them; on Linux they inherit the host thread's single-core pin and starve it (prompts read ~10x slower on a DGX Spark), and unpinned they take every core for the whole prompt.

This change makes both waits sleep:

- the ring wait uses C++20 `std::atomic::wait`, with `notify_all` in `issued_one()` and `finish()`;
- `dma_done` events are created with `cudaEventDisableTiming | cudaEventBlockingSync`;
- the HIP shim gets `cudaEventBlockingSync -> hipEventBlockingSync` (not tested on HIP).

No affinity changes. Measured on a DGX Spark (GB10, CUDA 13.0, Ubuntu 24.04) with IQ3_S, all experts in the GPU cache, prompts of 7,031 / 26,383 tokens:

| build | 8K read tok/s | 30K read tok/s | process CPU while reading | decode tok/s |
|---|---:|---:|---:|---:|
| before (44df8fd / same code as main) | 91 | 88 | ~100 (all on core 0) | 42-48 |
| stagers unpinned instead | 1,068-1,183 | 1,344-1,397 | 1,640-1,880 | 40-42 |
| **this PR** | **1,222-1,229** | **1,388-1,390** | **102** | **41-43** |

Output tokens are unchanged (greedy, same prompts). `git apply --check` passes on main (1735d64) and on PR #409's branch.

Measured with the help of Claude (Anthropic).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

No site

Links install, modelos, releases.