Pull requests / #792
#792 batch slots: the zero-doorbell graph (100% VRAM resident) rings no layer
closed · @0xPreDa · 0 comments · View on GitHub
BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux
Description
Fixes #776. ## Cause When a stage holds every one of its experts in VRAM, its verify window is recorded as the **zero-doorbell graph** (#646): `record_window` publishes no per-layer doorbell (`m_seq_` is never rung), and the only host wait left in the graph is layer 1's `wait_flag_ge(m_flag_, 1)`, before it copies the PLE rows. The solo window already handles this. In `Verifier::run` (`src/core/verify.cpp`, the `if (all_resident_ && !test_stall)` branch), the host gathers the PLE rows, raises `*flag = 1` once and skips the per-layer loop. The batch windows are recorded by the same `record_window`, but both batch host paths still wait for layer 0's ring: - `run_slot_rows` (without `--batch-groups`): `while (*seq < want)` never ends at `k = 0`. After 20 s it reports `verify batch: timed out at layer 0`, the engine exits and the server restarts it. - `batch_launch` / `batch_poll` (with `--batch-groups`): the same wait, in `batch_poll`. As a result, on any rig where a stage holds 100% of its experts, the **first** batch window fails - even with a single busy slot (`captured the batch window over slots 0`), with or without `--batch-groups`. The measurements in BATCHING.md were taken on cards that never reached 100% residency, which is why it did not show up there. #776 (2x RTX PRO 6000) and my 4x RTX 4090 split both reach it. ## Fix - `run_slot_rows`: when `all_resident_`, raise `h_flag_` to 1 once (`stage_batch` has already gathered the PLE rows before the launch) and skip the per-layer loop. The stream synchronization that follows waits for the window as before. - `batch_launch`: when `all_resident_`, raise `h_flag_` to 1 and set `b_k_ = b_steps_`, so that `batch_poll` only polls the stream. - `serve/server.py`, `_take_control`: `restart()` runs `__init__` again while requests are waiting for the control lines, which replaces `wait_lens` and resets `waiting`. The waiting request's `finally` block then raised `ValueError: list.remove(x): x not in list` and could drive `waiting` below zero; it showed up as `the engine reported an error: list.remove(x): x not in list` after each stall. The cleanup now tolerates a restart (identity check, `max(0, ...)`). Stages that are not fully resident (doorbell graph) are unchanged: the new branches only run when `all_resident_` is true. ## Testing Linux, 4x RTX 4090 (24 GB, no P2P), CUDA 13.0, sm_89, engine 0.1.39 with this patch, GSQ-RCO IQ3_S, `--layer-split 12,24,36 --trim-stage-weights --batch 4 --batch-groups 4`, prefill 2048, 262144-token context. All four stages report `100% VRAM resident: zero-doorbell graph`. Requests go through the HTTP server, with short prompts (0-12K tokens) and long ones (30K-100K tokens): | | Before | After | |---|---|---| | Batch windows | the engine exits at the first one (`verify batch: timed out at layer 0`), 9-10 restarts per run | no engine error, a single engine start for the whole run | | 4 concurrent requests, short prompts | 0 of 12 complete (the engine keeps restarting) | 12 of 12, 81 tok/s per request, first token in 1.1 s (median) | | 4 concurrent requests, 30K-100K prompts | - | 6 of 6, 50 tok/s per request, first token in 5.8 s (median) | | 1 request (solo path) | 178 tok/s | 177 tok/s | `serve.test_parallel`, `serve.test_lifecycle` and `serve.test_server` pass (158 tests). Not tested: Windows, HIP, and a split where some stages are fully resident and others are not.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.