Issues / #1603
#1603 parallel: 2 + two long-context sessions: the request hangs with running: 0 / queued: 0 and the #1317 frozen-engine watchdog never sees it
open · @ylzbj1-stack · 0 comentários · No GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quantsWindows
Descrição
# `parallel: 2` + two long-context sessions: the request hangs with `running: 0 / queued: 0` and the #1317 frozen-engine watchdog never sees it
**Engine:** 0.1.41 (`strata-windows-x64.zip`, `BUILD.json` → `version: 0.1.41`, `cuda 13.0`, `archs [75,86,89,120]`)
**Server:** `serve/server.py` from the same tag
**OS:** Windows 11, WDDM
## Related issues
- **#1317** (closed, fixed in 0.1.40.2 + 0.1.41): *an untimed read of the engine/vision pipe can hold the request turn forever*.
Here the request turn is **not** held: `/metrics` reports `running: 0`, `queued: 0`, `waiting: 0`
while the client is still waiting. So #1317's frozen-engine check (`ENGINE_STALL_S`, 90 s) has
**nothing to watch** and never fires — I grepped the whole engine log and the server stdout for
`has said nothing for` / `frozen` / `EngineSilent` and there is **no such line, ever**.
- **#481** (open): `serve engine deadlocks in SleepConditionVariableSRW during a long streaming reply`.
Same "lost step" family, but mine is reproducible by **two concurrent long sessions**.
- **#1341** (open): `a >100k-token prompt deadlocks the pipelined verify-window reader
(--pipeline-windows >= 1)`. **Not this one:** I do **not** pass `--pipeline-windows`, and my prompts
are ≈40 K, not >100 K.
## Environment
- 2× RTX 3080 20 GB (sm_86), `layer_split: auto` → CUDA0 layers 0–24, CUDA1 layers 25–47
- Xeon E5-2696 v3 (18 c / 36 t, **no AVX-512**), 128 GB DDR3-1600 quad channel
- Model: Swift-1.5 Qwen3.8-Flash-Next GSQ-RCO **IQ3_XXS**, native pack, shards 39.79 + 36.18 GB
- `parallel: 2`, `--max-context 262144`, `--kv-resident 32768`, `--kv int8`, `--spec 4`,
`--prefill auto`, `--expert-cache auto`, `--vram-reserve-mib 700`, `--vision`
- **No `--pipeline-windows`, no `--batch-groups`, no `--resident-experts`**
## Reproduction
Two **long-context** conversations (~40 K tokens each) active simultaneously on `parallel: 2`,
with at least one of them needing a chunked, partly-cold prefill. A later request (long or short)
then hangs in the client forever.
It is **not deterministic** — it needs the two long prompts to be in flight close together. Once it
happens it stays until the client cancels.
## Symptom — both states captured via `GET /metrics`
**Hung:**
```
live.state = "idle" live.running = 0
live.queued = 0 live.waiting = 0
live.slots = [ {slot 0, state "idle", held_tokens 885},
{slot 1, state "idle", held_tokens 38605} ]
netstat <port> : 0 × ESTABLISHED (only TIME_WAIT)
GPU power : 31.9 W / 7.1 W (idle; limits 320 W)
engine CPU : kernel 0.00 s + user 0.00 s over an 8 s sample → 0 % CPU
```
**Healthy, seconds later, after closing the second long session:**
```
live.state = "generating" live.running = 1 ESTABLISHED = 2
GPU0 105.5 W / GPU1 133.8 W
slot 0 held=43103 | slot 1 held=143
```
So the request is neither `running`, nor `queued`, nor `waiting`, and there is no TCP connection
held by it — yet the client is still on the "waiting for token generation" panel. Whatever booked it
in was released without ending the client's turn.
## Evidence
### 1. Both long requests admitted in the same second
`/metrics → requests` (server-side bookkeeping; every entry here ends `finish: "stop"`):
```
00:51:54 prompt=665 reused=0 out=221 44.1 s finish=stop
00:51:55 prompt=38385 reused=0 out=8473 129.8 s finish=stop <- admitted together
00:54:05 prompt=38620 reused=38380 out=4230 43.9 s finish=stop
00:54:50 prompt=39086 reused=38615 out=173 3.1 s finish=stop
00:54:54 prompt=39300 reused=39081 out=78 1.9 s finish=stop
00:54:57 prompt=39452 reused=39295 out=128 2.2 s finish=stop
00:58:04 prompt=40024 reused=39447 out=893 16.3 s finish=stop
```
The server believes it completed everything. The client-side turn never ended.
### 2. BYIELD fires repeatedly — the two slots keep interrupting each other's reads
```
strata batch: the prompt was read in 6 parts, the slots decoding 21138 ms between them
strata batch: the prompt was read in 3 parts, the slots decoding 10754 ms between them
strata batch: the prompt was read in 2 parts, the slots decoding 2206 ms between them
strata batch: the prompt was read in 3 parts, the slots decoding 4180 ms between them
strata batch: the prompt was read in 3 parts, the slots decoding 4105 ms between them
```
### 3. Control experiment
- One long session + an empty second slot → **stable**, 100.3 / 102.4 / 108.5 tok/s,
expert-cache hit 95–98 %, KV streaming 99.7 % of block reads hitting VRAM, drafts ≈ 72–80 %.
- One long session + a *short* request → fine (a short prompt does not need to yield).
- **Two long sessions → hang.** Closing one of them resolves it immediately.
### 4. The transport is not the problem
A streaming request to the same server completes:
`chunks=77`, `finish_reason=stop`, `data: [DONE]` received (`server.py:4985`), 3.5 s.
### 5. Nothing in the log about the hang
Neither the engine log nor the server's stdout has any line for the stalled request — no
`engine_silence_s` / `STRATA_ENGINE_STALL_S` notice, no engine restart, no error. The request simply
never appears.
## Where I would look
The admission path in `server.py` around the `d9621d1` comment (0.1.41):
```python
# With none free the control lines are given back while waiting: a slot can be held by a read
# that gave way (BYIELD) and needs them to go on, so waiting with them is a deadlock - long,
# long, short at 0.25 s steps under parallel: 2 (ENGINE_REVIEW finding 1, reproduced in 0.1.41).
while True:
with self.slot_cv:
slot = self.pick_slot(prompt)
if slot is not None:
self.slot_busy[slot] = True
break
if holding:
self.ctl.release()
holding = False
...
if not holding:
ok = yield from self._take_control(cancel, len(prompt))
```
`d9621d1` handles "waiting for a slot". Mine is the case where a **yielded read** (the
`if self._yielded is not None:` branch just below, and the `BYIELD` sends at `_send(f"BYIELD {slot}")`)
and a **waiting admission** contend for the same control lines while a third request is also waiting.
The end state (`running: 0` **and** `queued: 0` **and** `waiting: 0` with the client still waiting)
suggests the request is dropped from the bookkeeping rather than merely blocked — which is also why
`ENGINE_STALL_S` cannot catch it: the frozen-engine check only runs while a request is outstanding.
## What I can provide
- Full engine log around the incident
- `/metrics` snapshots in both states (hung and healthy)
- The config JSON
- I can run an instrumented build, `STRATA_DBG_FEAT`, `ROUTE_OVERRIDE`, or any env flags you want
## Workaround (for other users hitting this)
Do not run two long-context sessions concurrently on `parallel: 2`. Short requests are safe to run
concurrently — they do not trigger a BYIELD yield. A single long session is stable at 100+ tok/s.
No site
Links install, modelos, releases.