Issues / #1012
#1012 0.1.39 serve: after an engine restart, waiting requests hang forever or fail with "list.remove(x): x not in list"
closed · @Jackwwg83 · 1 コメント · GitHub で見る
BenchmarksServer & APINVIDIA / CUDAModels & quantsLinux
本文
**Summary:** after the server restarts a failed engine, requests that were already waiting for the control lines never finish. Others fail with `list.remove(x): x not in list`. In our run, 5 requests stayed at "reading the prompt" for over 90 minutes. `/health` said `loaded: true`, and the GPU was idle. **Version:** Strata v0.1.39 (`6f32ec0`), Linux, RTX PRO 5000 Blackwell 48 GB, IQ3_S, `"parallel": 8`. **What triggers it:** any engine exit while requests wait. In our case the trigger was the batch-window OOM from #997 (`verify: batch instantiate: out of memory`) with 10 concurrent users. **Server log** ``` [strata] the engine reported an error: verify: batch instantiate: out of memory [strata] the engine had stopped (exit code 1); starting it again (a minute or two) ... [strata] starting the engine: reading the model's weights ... [strata] reading the prompt: 34,465 tokens, 10 s so far ... [strata] the engine is running again [strata] the engine reported an error: list.remove(x): x not in list [strata] done: 0 tokens in 14 s (0.0 tok/s) (error, cancel=False) ... [strata] the engine reported an error: list.remove(x): x not in list ... [strata] reading the prompt: 30,568 tokens, 5531 s so far [strata] reading the prompt: 29,188 tokens, 5521 s so far [strata] reading the prompt: 34,712 tokens, 5531 s so far [strata] reading the prompt: 45,184 tokens, 5531 s so far [strata] reading the prompt: 28,099 tokens, 5521 s so far ``` The last five lines repeat every 10 s until the clients disconnect. **Cause (from reading `serve/server.py` at v0.1.39)** - `StrataEngine.restart()` (line 577) restarts the engine by calling `self.__init__(*self.spawn)` (line 590) on the live object. - `__init__` creates the admission state again (lines 507-512): `slot_cv`, `waiting`, `wait_lens`, `ctl_epoch` and the `ctl` lock. - A request thread that is already in `_take_control()` (line 815) still waits on the old `ctl` lock and the old `slot_cv`. Nothing notifies or releases them again, so these requests hang (the 10 s heartbeats above). - A thread that leaves `_take_control()` runs `self.wait_lens.remove(entry)` (line 840) on the new, empty list. This raises `ValueError: list.remove(x): x not in list`, and the request fails. `self.waiting` also ends up below zero. **Possible fix:** keep the admission state (locks, condition, wait list and counters) outside the part of `__init__` that `restart()` runs again. Or, in `restart()`, end every waiting request with an error before the engine starts again, so that clients can retry.
関連リンク
インストール・モデル・リリースへの站内リンク。