Issues / #1118

#1118 Batch slots: the engine stalls ("no progress for 60 s … reading the prompt (batched)") when a long prompt is read while another slot decodes; STRATA_BATCH_DECODE_SHARE=0 avoids it

closed · @uncle-daddy-jp · 3 comments · View on GitHub

BenchmarksServer & APINVIDIA / CUDAModels & quantsWindowsLinux

Description

**Summary**

With `"parallel": 3` (`--batch 3`), the engine stalls and then kills itself ("no progress for 60 s during a request (reading the prompt (batched), done up to token N)") when a long prompt is being read while another slot is decoding. It reproduces every time with real agent traffic (a ~6K prompt generating in slot 0 while a ~21K prompt is read), and in a synthetic load of 3 concurrent requests (6.4K / 21.5K / 8.2K tokens). It does not reproduce with short prompts. `STRATA_BATCH_DECODE_SHARE=0` makes it go away; `"parallel": 1` of course too. `--prefill 2048`, `--batch 2`, and turning the conversation cache off do not help. Two server-side follow-ups are listed below (a 400 after an engine restart, and requests that stay open for ~6 minutes after the engine died).

**Environment**

- NVIDIA DGX Spark (GB10, aarch64, 128 GB unified memory), driver 580.142, CUDA 13.0, Ubuntu 24.04.
- Strata main 1735d64 (the one before today's history rewrite) + PR #1057/#1101 (stager sleep) + the aarch64 port from PR #409 (44df8fd) cherry-picked. Engine reports 0.1.40. The batched prompt path is upstream code; the port only touches build, UMA sizing and ARM CPU kernels, so I expect this to reproduce on x86 Linux with `--mmap-experts` as well, but I have not tried that.
- Model: Qwen3.8-Flash-Next IQ3_S (GSQ-RCO), `--mmap-experts --expert-cache auto --prefill auto --spec 4 --mtp … --kv int8 --max-context 131072 --vision`, all 24,576 experts in the GPU cache, control vector on (also reproduced with it off).
- Melody/other services were running too, but RAM was not the problem: 30 GB free at the stall, swap 274 MB, zero OOM kills, all 19 pool workers sleeping.

**Repro**

1. `"parallel": 3` in `strata-<model>.json` (plus `--conversation-cache-mib 12288 --conversation-cache-slots 12`, but it also reproduces without the conversation cache).
2. Send three chat completions at once: A = 6.4K prompt, thinking on (`reasoning_effort: medium`, `reasoning_budget_tokens: 4096`, max_tokens 1500); B = 21.5K prompt, thinking on, two tools; C = 8.2K prompt, thinking off, max_tokens 300.
3. A gets admitted and decodes in slot 0. While B's prompt is read, the read stops at a chunk boundary and 60 s later the engine exits with code -6.

Engine log at the stall (from `strata-<model>.log`):

```
strata batch: BYIELD 1 not taken at 14641 of 21543
strata verify: layers [0, 48) have experts out of VRAM (a prompt loan, a shrink or a swap): the doorbell graph runs the windows
strata verify: captured the batch window over slots 0
strata serve: no progress for 60 s during a request (reading the prompt (batched), done up to token 14641) - stopping the engine so the server starts it again (issue #29)
```

The stall report: "0 of 0 jobs claimed, 19 of 19 workers parked", "the GPU rang 1; flags: served 1, plan (A) 0, copies (B) 0", "released the verify window's GPU waits (#267): the GPU finished in 27 ms". In production the same four lines appeared four times in ten minutes (done up to token 8192 each time, i.e. the first chunk boundary), each followed by a 30-60 s engine restart.

What it looks like from the outside: between prompt chunks the server lets the decoding slot run (`STRATA_BATCH_DECODE_SHARE`, default 0.5); because the prompt path has borrowed cache slots, that decode window runs as the doorbell graph; the window is captured and then nothing progresses.

**What was tried (same 3-request load, 3 rounds each)**

| config | result |
|---|---|
| `--batch 3`, prefill auto (8192) | stall + exit -6 in round 1 |
| `--batch 3`, `--prefill 2048` | stall at token 2048, exit -6 |
| `--batch 2`, conversation cache 4096 MiB | stall, exit -6 |
| `--batch 3`, conversation cache off | stall, exit -6 |
| `--batch 3` + `STRATA_BATCH_DECODE_SHARE=0` | no stall, 2 of 2 runs; the other slots' decode drops to 5-11 tok/s while a prompt is read (expected) |
| `--batch 1` | no stall; 104 s for the 3 rounds vs 117-124 s with SHARE=0 |

**Two server-side follow-ups seen during the restarts**

1. `list.remove(x): x not in list` → `400 invalid request`. `serve/server.py` recreates `wait_lens` / `slot_cv` when it restarts the engine (around line 660); a request that was queued across the restart then hits `self.wait_lens.remove(entry)` in its `finally` and the client gets a 400 for a request it never got to send.
2. After the engine exits, the requests that were reading or waiting keep printing "reading the prompt: 14,641 of 21,543 tokens, 380 s so far" until `engine_silence_s` (300 s) closes them with an error, ~384 s in total. Clients with a 120 s timeout have long given up; failing them at once when the engine dies would be kinder.

**Workaround in use**

`"parallel": 1` with the conversation cache (`--conversation-cache-slots 12`). Stable so far under the same traffic.

Logs (engine log, server stdout, stall reports, the load script) are available if useful. Measured with the help of Claude (Anthropic).

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.