Issues / #1619
#1619 serve: a request waiting for the single request turn sends no bytes, so clients with a stream idle timeout abort it (and the retry re-reads the prompt)
open · @taxah92 · 0 コメント · GitHub で見る
Server & APINVIDIA / CUDAModels & quants
本文
## Summary With the default single-turn server (`parallel` / `--batch` off) a request that is waiting for the request turn receives **no bytes at all** — not even SSE comments. The 10 s heartbeat exists only for the request that already holds the turn (`serve/server.py:1472`, `1145`), and the FIFO is taken before the first `yield` of the response (`serve/server.py:3466`, first `yield` at `3489`). A client with a stream idle timeout therefore aborts a perfectly healthy queued request; ours retries it, and each retry starts a new prompt read, so the queue grows and the loop repeats. ## Environment - Strata v0.1.41, engine 0.1.41, built from source (sm_70), Ubuntu 26.04, Tesla V100 SXM2 16 GB - Qwen3.8-Flash-Next IQ3_S, `--max-context 1048576` (YaRN), `--kv k8v4 --kv-resident 32768`, `--spec 3`, `--prefill auto:16384`, conversation cache 24576 MiB / 3 slots, `parallel` not set (one request at a time) - Client: OpenAI-compatible over SSE with a 300 s stream idle timeout ## What we see `journalctl -u strata-0141` (three agent contexts sharing one server): ``` 01:05:41 done: 25406 tokens in 529 s (stop) <- one request generated for 8.8 minutes 01:05:42 done: 0 tokens in 1 s (cancel) <- two requests that had waited ~5 minutes 01:05:44 done: 0 tokens in 2 s (cancel) 01:06:53 done: 0 tokens in 69 s (cancel) <- their retries 01:07:03 done: 0 tokens in 10 s (cancel) 01:08:52 done: 585 tokens in 109 s (stop) <- only now service resumes ``` The two aborted requests show 1-2 s of engine time: they were ended right after they finally got the turn, so they had spent the whole time queued. The engine log for the retries: ``` strata serve: prompt 141160 tokens = 12166 reused + 91392 of 128994 read in 68815 ms ... 0 generated ``` ## Why it matters The abort also costs a conversation-cache slot (the restored entry is taken out of the cache and a cancelled turn never parks it — see the companion report), so the retry re-reads everything after the root checkpoint. We counted 28 such events in one day: ~740k tokens ≈ 9-10 minutes of pure re-prefill. ## Suggested fix Send the same heartbeat while a request waits for the FIFO as while it waits for the engine: SSE headers plus a `: keepalive` comment every ~10 s around the `self.fifo` acquisition (the heartbeat path already exists — `serve/server.py:1145`, `1472`). SSE comments are the standard liveness signal and would also notice a client that went away while queued. A 300 s idle timeout is not an unusual client setting, and on a 16 GB card the queue can be 8+ minutes long.
関連リンク
インストール・モデル・リリースへの站内リンク。