Pull requests / #1025
#1025 fix: unblock all-resident batch windows (batch verify waited for a doorbell the graph never publishes)
closed · @win10ogod · 0 Kommentare · Auf GitHub
BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsWindows
Beschreibung
The all-resident ("zero-doorbell") verify graph deliberately omits the per-layer doorbell publications, but the two batch host paths still waited for the layer-0 doorbell. A batch window therefore stalled for the full 20 s, released the GPU waits, and the engine exited 1:
```
verify batch: timed out at layer 0; its GPU waits were released and the GPU finished (#267)
the engine stopped unexpectedly (exit code 1)
```
**Reproduce:** two or more concurrent requests through the server on an all-resident pack. It fails at 2, 4 and 8 concurrent alike; one request at a time is unaffected. The server switches from the solo path to the slots as soon as a second request arrives, which is why a batch window can initially hold only slot 0.
**Mechanism:** `record_window()` skips every doorbell publication when `all_resident_`, so `h_seq_` stays 0 and nothing ever publishes `seq >= 1`. Meanwhile `stage_batch()` resets `h_flag_` to 0 while the captured graph still contains `wait_flag_ge(m_flag_, 1)` guarding the layer-1 PLE input. The host waits for a signal the graph does not produce; after the 20 s timeout the release lets the PLE wait proceed, which is why the message says the GPU finished. The solo path already handles this at the `all_resident_` branch; the batch paths did not.
## Fix (`src/core/verify.cpp`, 14 insertions / 3 deletions)
- `stage_batch()`: for all-resident windows, publish `h_flag_ = 1` after the fences, using the same ordering the solo path uses.
- `run_slot_rows()`: expect zero doorbell steps when `all_resident_`, so the existing graph-completion checking, sampling and output collection run instead of the doorbell loop.
- `batch_launch()`: zero expected doorbell steps for an all-resident pipeline stage, so `batch_poll()` proceeds to its completion checks.
- `release_gpu_waits()`: a CUDA error is no longer reported as drained - only `cudaSuccess` counts. Previously any result other than `cudaErrorNotReady` was treated as a finished GPU, which made "the GPU finished" unreliable.
The zero-doorbell optimization is kept; batching is not disabled, and no timeout was raised.
## Verification on this machine
RTX PRO 6000 Blackwell (sm_120), driver 616.92, CUDA 13.3, Release build, `STRATA_ENABLE_CUDA=ON`, architecture 120. Isolated server on port 18081 with `parallel: 8` and **no** `STRATA_VERIFY_ALL_RESIDENT` workaround; the engine log confirms the zero-doorbell path was active (`100% VRAM resident: zero-doorbell graph`).
| Test | Before | After |
| --- | --- | --- |
| 8 concurrent x 3 rounds | 0/8 every round, engine exit 1 | **24/24**, no timeout, no engine restart |
| 2 and 4 concurrent | 0/2, 0/4 | pass |
| Single request (greedy) | - | 121.5 tok/s, drafts 11/30 |
| Repeated greedy prompt | - | identical output |
| Long prompt (9,126 tokens) beside 3 active decoders | - | completed in 3.7 s, 3/3 decoders unaffected |
Host-side checks: `tools/test_setup_parallel.py`, `tools/test_setup_choices.py`, `tools/test_setup_amd.py` - 66/66 passed; `git diff --check` clean.
## Not covered
No multi-GPU or pipeline verification (single card here). Phase-1 analysis also flagged an adjacent issue that this patch does **not** address: prefill can temporarily mark experts non-resident while slots decode, whereas `all_resident_` is latched at initialization. That needs its own investigation; the long-prompt test above passed but is not a proof for it.
Mehr auf der Site
Links zu Install, Modellen, Releases.