Issues / #1059

#1059 serve: a request the engine rejects with ERR waits the full 300 s before failing (STOP + _drain_control("DONE") after an ERR), holding the control lock

closed · @cardosofelipe · 1 comments · View on GitHub

BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quantsLinux

Description

**Summary:** when the engine rejects a request with an `ERR …` line, the server waits the full 300 s before failing it, and appears to hold the engine's control lock while it waits. A request the engine refuses immediately takes 5 minutes to fail.

**Version:** v0.1.39 (`6f32ec07`), observed on Linux, RTX 3090. The same code path is unchanged in v0.1.40 (`serve/server.py`, `_drain_control` line 937 and the `finally` block around line 1225; `generate.cpp` line 8081).

**What happens**
1. The engine refuses a request up front: it prints `ERR …` and `continue`s back to its input loop. It never prints `DONE` for that request.
2. The server's generator raises `ValueError` from the `ERR` line (`[strata] the engine reported an error: …`), as expected.
3. The `finally` block then runs, with `phase == "solo"`: it sends `STOP` and calls `_drain_control("DONE")`.
4. The engine is idle, so no `DONE` (or further `ERR`) line arrives. `_drain_control` waits out its hard-coded `timeout=300.0`.
5. The request ends `done: 0 tokens in 300 s (0.0 tok/s) (error, cancel=False)`. The client sees nothing until then; a client with a shorter idle timeout gives up first.

**Reproduction** (any up-front engine rejection works; this is how we hit it)
- A config with a `"vision"` section but without `--vision` in the engine args (our mistake; setup adds both).
- Send a chat completion containing an `image_url` part.
- Log: `[strata] the engine reported an error: this engine was started without --vision`, then `done: 0 tokens in 300 s (0.0 tok/s) (error, cancel=False)`. Six image requests in a row took 6 × 300 s; GPU idle at 0 % throughout.

Other `ERR`s on the same pre-generation path should behave the same (from reading the code, not tested): `ERR bad request`, `ERR prompt (…) + max_new (…) exceeds the context`, `ERR a token id is outside the vocabulary`, `ERR the K/V cannot grow to this prompt: no VRAM is left`. The last one seems the most likely to occur in normal use with `"parallel"` and long contexts.

**Impact beyond the one request (from reading the code, not measured):** the drain runs before `self.ctl.release()`, so other requests appear to be blocked for the same 300 s.

**Possible fix:** when the request ended because of an engine `ERR` line, the engine has already finished with it, so skip the `STOP` + `_drain_control` (or bound that drain to a few seconds). The engine already discards a stale `STOP` between requests (`stop_req.store(false)` before a new request), so not sending one is safe.

Happy to test a patch on this setup.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.