Pull requests / #845
#845 verify: batch windows honor the all-resident stage (no host doorbells)
closed · @Yoshiki-Matsuda · 0 comentários · No GitHub
Server & APIMulti-GPUNVIDIA / CUDAWindows
Descrição
Combining `--layer-split` with `--batch` (parallel >= 2) killed the engine: "verify batch: layer 36 never rang (graph finished)", exit code 1. Cause: #646's all-resident optimization plans on the device and raises no host doorbells, and the solo path (`run`) branches around its per-layer host service for such a stage. `run_slot_rows` did not: it spun on the doorbell ring for every layer of the resident stage (the 36-47 stage on the second GPU), waiting on a ring nothing rings, until the 20 s timeout ended the engine. Solo-only sessions never hit it; the batch window is the first path through `run_slot_rows`. Fix: - `run_slot_rows`: the same all-resident branch `run()` makes - raise the PLE flag the graph's first wait reads, sync the stream, chain to the next stage or sample the rows. - `batch_poll`: skip the per-layer service loop for an all-resident stage, and raise the PLE flag there too (the mirror case: stage 0 resident, stage 1 not). - `collect_profile()` on the fast path so profiling keeps its numbers. Measured on RTX 5090 + RTX 3090, split 36, batch 2-3: 3 concurrent requests all complete, zero "never rang", batch windows run at ~1.9 rows avg. `python -m serve.test_parallel`: 11 tests pass.
No site
Links install, modelos, releases.