Pull requests / #761

#761 stage no longer waits for the whole split below it

closed · @gopinath87607 · 0 评论 · 在 GitHub 查看

BenchmarksMulti-GPUModels & quants

描述

A 4-way layer split's stages took turns.  72% of one-second samples during a 120,000-token prompt had exactly one of the four cards busy and never three or four, for 1,090 tok/s.

The cause was not the hand-off buffer count.  `Prefill::run` ended with `if (next_run.valid() && !next_run.get())`, so a stage's `run` did not return until its successor's had, and its successor's not until *its* successor's - every stage inherited the whole downstream chain's latency per chunk.  Measured on this rig (2x3060 + 2x5060, IQ3_S, --layer-split auto, chunk 10,240): dev0 spent 6,253 ms of every 8,137 ms chunk inside that one `get()`, while its own GPU work was 2,234 ms.

The chunks below are now owned by a per-stage deque of futures, reaped one-for-one before the next hand-off, so a stage returns while the chunks below it are still running.  What remains is that the OUTERMOST stage must not return with chunks still in flight, or the caller tears the split down under them: a guard on `run` covers every return path, and it calls `Prefill::finish` (which a caller can also call to join the split early).

  1,089.7 / 1,090.1 tok/s  ->  2,210.6 / 2,227.5 tok/s         2.04x

Same binary, one prompt, two runs per arm; the two arms of a pair differ by 0.2% and 0.8%. Cards busy at the same moment went from 72%/21%/0% (one/two/three-or-more) to 13%/19%/54%, and the wait at the hand-off from 6,253 to 131 ms on dev0 (dev1 3,289 -> 56, dev2 1,680 -> 55). Each stage's own work rose under the overlap (dev0 2,234 -> 3,246 ms), so the gain is smaller than the wait it removes.  The residual dumps of all four arms are byte-identical (md5 af2c193e86892f6f31abe160cc304d07), which is the point: this reorders waits, never arithmetic.

A stage owns ONE set of chunk buffers (m.R, the staging ring, the GEMM scratch), so two of its chunks must never be in flight at once - the reap is one-for-one, and STRATA_PF_DEPTH above 1 is refused rather than silently racing them through the same buffers.  STRATA_PF_DEPTH=0 restores the old join, so the A/B is one binary apart.  The hand-off buffer goes from two to one (the copy now happens after the join instead of before), which halves its pinned host memory, 800 -> 400 MiB at chunk 10,240.  STRATA_PREFILL_TIMING also reports each stage's wall time and how much of it went into that wait.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。