Issues / #1012
#1012 0.1.39 serve: after an engine restart, waiting requests hang forever or fail with "list.remove(x): x not in list"
closed · @Jackwwg83 · 1 comments · View on GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quantsLinux
Description
**Summary:** after the server restarts a failed engine, requests that were already waiting for the control lines never finish. Others fail with `list.remove(x): x not in list`. In our run, 5 requests stayed at "reading the prompt" for over 90 minutes. `/health` said `loaded: true`, and the GPU was idle. **Version:** Strata v0.1.39 (`6f32ec0`), Linux, RTX PRO 5000 Blackwell 48 GB, IQ3_S, `"parallel": 8`. **What triggers it:** any engine exit while requests wait. In our case the trigger was the batch-window OOM from #997 (`verify: batch instantiate: out of memory`) with 10 concurrent users. **Server log** ``` [strata] the engine reported an error: verify: batch instantiate: out of memory [strata] the engine had stopped (exit code 1); starting it again (a minute or two) ... [strata] starting the engine: reading the model's weights ... [strata] reading the prompt: 34,465 tokens, 10 s so far ... [strata] the engine is running again [strata] the engine reported an error: list.remove(x): x not in list [strata] done: 0 tokens in 14 s (0.0 tok/s) (error, cancel=False) ... [strata] the engine reported an error: list.remove(x): x not in list ... [strata] reading the prompt: 30,568 tokens, 5531 s so far [strata] reading the prompt: 29,188 tokens, 5521 s so far [strata] reading the prompt: 34,712 tokens, 5531 s so far [strata] reading the prompt: 45,184 tokens, 5531 s so far [strata] reading the prompt: 28,099 tokens, 5521 s so far ``` The last five lines repeat every 10 s until the clients disconnect. **Cause (from reading `serve/server.py` at v0.1.39)** - `StrataEngine.restart()` (line 577) restarts the engine by calling `self.__init__(*self.spawn)` (line 590) on the live object. - `__init__` creates the admission state again (lines 507-512): `slot_cv`, `waiting`, `wait_lens`, `ctl_epoch` and the `ctl` lock. - A request thread that is already in `_take_control()` (line 815) still waits on the old `ctl` lock and the old `slot_cv`. Nothing notifies or releases them again, so these requests hang (the 10 s heartbeats above). - A thread that leaves `_take_control()` runs `self.wait_lens.remove(entry)` (line 840) on the new, empty list. This raises `ValueError: list.remove(x): x not in list`, and the request fails. `self.waiting` also ends up below zero. **Possible fix:** keep the admission state (locks, condition, wait list and counters) outside the part of `__init__` that `restart()` runs again. Or, in `restart()`, end every waiting request with an error before the engine starts again, so that clients can retry.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.