Issues / #997

#997 0.1.39: engine exits (code 1) with "parallel": 4 on UD-IQ4_XS, RTX 3090 24 GB, three concurrent 8K prompts (reproducible; IQ3_S unaffected)

open · @cardosofelipe · 5 comments · View on GitHub

BenchmarksServer & APINVIDIA / CUDAModels & quantsLinux

Description

**Summary:** with `"parallel": 4`, UD-IQ4_XS on a 24 GB RTX 3090 exits with code 1 (no error line) when three 8K-token prompts are read concurrently. Reproduced 2/2, at the same point both times. IQ3_S with the identical test and `"parallel": 4` never failed (2 × 21/21 requests).

**Environment**
- Strata v0.1.39 (`6f32ec07`), Linux build (Fedora 44, kernel 6.19.10), CUDA 13.3, driver 595.80
- RTX 3090 24 GB, Ryzen 9 5950X (AVX2, no AVX-512), 128 GB DDR4 dual-channel
- Model: Qwen3.8-Flash-Next UD-IQ4_XS (Unsloth, 3-part GGUF), native pack, MTP draft layer
- Engine args: `--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.70 --mtp <rt> --max-context 32768 --kv int8 --resident-budget-gib 55 --pcie-frac 0.20` (the second run; the first used defaults `--spec-min-p 0.5`, no `--pcie-frac`, same result), `"parallel": 4`
- Startup: `--batch: 4 slot sessions on CUDA0 (0.56 GiB each); 15.89 GiB free`, then `360 MiB of VRAM free with everything loaded`
- Host RAM at the time: ~72 GiB available, no swap used; no systemd OOM kill; unit memory peak 48.2 GB

**Reproduction**
OpenAI streaming chat completions, unique ~8,150-token prompts, `max_tokens` 512, reasoning on. Concurrency ladder 1 → 2 → 3, three waves per level:
- Level 1 (3/3) and level 2 (6/6) passed: ~57 and ~28 tok/s per request.
- Level 3: waves 1 and 2 passed (~20 tok/s each). On wave 3 the three slots were reading their prompts (`slot 1 takes 8151 tokens`, `slot 2 takes 8152 tokens`, prompt reads at ~1,860–1,885 tok/s). Each had generated 2–4 tokens when the engine exited.

**Log at failure**
```
strata serve: prompt 8152 tokens = 0 reused + 8152 read in 4324 ms (1885.1 tok/s), 1 generated in 25 ms (39.6 tok/s), drafts accepted 0 of 0, 1 checkpoints
strata serve: decode expert cache hit rate: 57.9% (271 hits / 468 lookups); 12 more read by the GPU over PCIe or from another GPU (2.5% of all 480 routed)
strata serve: resident RAM: 41.75 GiB of experts in RAM, 13981 exchanged with the VRAM tier, 27846 blob reads from the file
strata serve: expert tiers: GPU 271 hits this request; since the start RAM 1232526 blobs, files 27846 blobs 149546.1 MB read (the GGUF in place)
[strata] thinking: 4 of max 512 tokens, 0.3 tok/s, 19 s
[strata] the engine stopped unexpectedly (exit code 1). The engine exited (code 1). Its last log line: strata serve: expert tiers: GPU 271 hits this request; ...
```
All three in-flight requests failed (`done: 1/2/4 tokens in 23 s (error, cancel=False)`). The engine log has no traceback, CUDA error or out-of-memory message.

**Possibly related:** #511 / #535 (stalls with very little free VRAM). Here 360 MiB was free after loading. I haven't tried a larger `--vram-reserve-mib` yet; that would be my next test if useful.

Happy to provide full engine/server logs or run a debug build.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.