Pull requests / #845

#845 verify: batch windows honor the all-resident stage (no host doorbells)

closed · @Yoshiki-Matsuda · 0 コメント · GitHub で見る

Server & APIMulti-GPUNVIDIA / CUDAWindows

本文

Combining `--layer-split` with `--batch` (parallel >= 2) killed the engine: "verify batch: layer 36 never rang (graph finished)", exit code 1.

Cause: #646's all-resident optimization plans on the device and raises no host doorbells, and the solo path (`run`) branches around its per-layer host service for such a stage. `run_slot_rows` did not: it spun on the doorbell ring for every layer of the resident stage (the 36-47 stage on the second GPU), waiting on a ring nothing rings, until the 20 s timeout ended the engine. Solo-only sessions never hit it; the batch window is the first path through `run_slot_rows`.

Fix:
- `run_slot_rows`: the same all-resident branch `run()` makes - raise the PLE flag the graph's first wait reads, sync the stream, chain to the next stage or sample the rows.
- `batch_poll`: skip the per-layer service loop for an all-resident stage, and raise the PLE flag there too (the mirror case: stage 0 resident, stage 1 not).
- `collect_profile()` on the fast path so profiling keeps its numbers.

Measured on RTX 5090 + RTX 3090, split 36, batch 2-3: 3 concurrent requests all complete, zero "never rang", batch windows run at ~1.9 rows avg. `python -m serve.test_parallel`: 11 tests pass.

関連リンク

インストール・モデル・リリースへの站内リンク。