Issues / #1012

#1012 0.1.39 serve: after an engine restart, waiting requests hang forever or fail with "list.remove(x): x not in list"

closed · @Jackwwg83 · 1 评论 · 在 GitHub 查看

BenchmarksServer & APINVIDIA / CUDAModels & quantsLinux

描述

**Summary:** after the server restarts a failed engine, requests that were already waiting for the control lines never finish. Others fail with `list.remove(x): x not in list`. In our run, 5 requests stayed at "reading the prompt" for over 90 minutes. `/health` said `loaded: true`, and the GPU was idle.

**Version:** Strata v0.1.39 (`6f32ec0`), Linux, RTX PRO 5000 Blackwell 48 GB, IQ3_S, `"parallel": 8`.

**What triggers it:** any engine exit while requests wait. In our case the trigger was the batch-window OOM from #997 (`verify: batch instantiate: out of memory`) with 10 concurrent users.

**Server log**
```
[strata] the engine reported an error: verify: batch instantiate: out of memory
[strata] the engine had stopped (exit code 1); starting it again (a minute or two) ...
[strata] starting the engine: reading the model's weights ...
[strata] reading the prompt: 34,465 tokens, 10 s so far
...
[strata] the engine is running again
[strata] the engine reported an error: list.remove(x): x not in list
[strata] done: 0 tokens in 14 s (0.0 tok/s) (error, cancel=False)
...
[strata] the engine reported an error: list.remove(x): x not in list
...
[strata] reading the prompt: 30,568 tokens, 5531 s so far
[strata] reading the prompt: 29,188 tokens, 5521 s so far
[strata] reading the prompt: 34,712 tokens, 5531 s so far
[strata] reading the prompt: 45,184 tokens, 5531 s so far
[strata] reading the prompt: 28,099 tokens, 5521 s so far
```
The last five lines repeat every 10 s until the clients disconnect.

**Cause (from reading `serve/server.py` at v0.1.39)**
- `StrataEngine.restart()` (line 577) restarts the engine by calling `self.__init__(*self.spawn)` (line 590) on the live object.
- `__init__` creates the admission state again (lines 507-512): `slot_cv`, `waiting`, `wait_lens`, `ctl_epoch` and the `ctl` lock.
- A request thread that is already in `_take_control()` (line 815) still waits on the old `ctl` lock and the old `slot_cv`. Nothing notifies or releases them again, so these requests hang (the 10 s heartbeats above).
- A thread that leaves `_take_control()` runs `self.wait_lens.remove(entry)` (line 840) on the new, empty list. This raises `ValueError: list.remove(x): x not in list`, and the request fails. `self.waiting` also ends up below zero.

**Possible fix:** keep the admission state (locks, condition, wait list and counters) outside the part of `__init__` that `restart()` runs again. Or, in `restart()`, end every waiting request with an error before the engine starts again, so that clients can retry.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。