Issues / #1577
#1577 Strata 0.1.40 on a Tesla T4 (sm_75): every request stalls in the first verify window, the watchdog stops the engine
open · @Madingdang · 0 comments · View on GitHub
BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindows
Description
[strata-iq2_xs.log](https://github.com/user-attachments/files/33210361/strata-iq2_xs.log)
[strata-stall-22088.dmp](https://github.com/user-attachments/files/33210374/strata-stall-22088.dmp)
## Summary
On a Tesla T4 (Turing, compute capability 7.5, driver 616.92) the engine loads the model and the server reports
`loaded: true`, but the **first request never produces a token**: the host waits in a verify window, the 60 s
watchdog stops the engine (issue #29 behaviour) and the server answers `HTTP 503`. Reproduced with two different
models (IQ3_XXS and IQ2_XS), on two different drives, with 15 GB RAM still free, and with 11 different
configurations/engines settings. The GPU is idle (0% utilisation) while the host waits, and a stall report +
`strata-stall-<pid>.dmp` is written every time.
## Environment
| | |
| --- | --- |
| OS | Windows 11 专业版, 10.0.22621 |
| GPU | NVIDIA Tesla T4, 16 GB VRAM, compute capability 7.5, driver 616.92 |
| CPU | AMD Ryzen 7 5700G (AVX2, no AVX-512) |
| RAM | 60 GB |
| Engine | 0.1.40 ready-made, CUDA 13.0, `archs [75, 86, 89, 120]`, `ptx: true` (`engine/BUILD.json`) |
| Storage | model on NVMe (`M2`, 2.46 GiB/s during load) |
Models used for reproduction (both are the plain community shards, both load fine in llama.cpp b11457):
* `Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-00001/00002-of-00002.gguf` (76 GB)
* `Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS-00001/00002-of-00002.gguf` (68 GB)
## What happens
1. Start: `run-iq2_xs.bat` (engine 0.1.40). Startup is fine:
```
strata generate: loaded 33.02 GiB at 2.46 GiB/s
strata generate: expert cache 7525 slots, 10.08 GiB of VRAM; policy is PROFILE, ranked by routing frequency
strata generate: R4 hit path ON - resident experts are computed on the GPU
strata serve: prompt chunk auto: 8192 tokens, a 384-slot ring
strata verify: captured the 4-token window (upload no error, sync no error)
strata serve: 474 MiB of VRAM free with everything loaded
```
`/health` → `{"status": "ok", ..., "model": "qwen3.8-flash-next-iq2_xs", "loaded": true}`.
2. Send one request (`/v1/chat/completions`, 17-token prompt, `reasoning_effort: none`). Nothing comes back.
After 60 s the server prints:
```
strata serve: no progress for 60 s during a request (reading the prompt (verify windows), from token 0)
- stopping the engine so the server starts it again (issue #29)
strata serve: stall report (engine 0.1.40): stage "reading the prompt (verify windows), from token 0" for 60 s
expert pool: epoch 0, batch epoch 0: 0 of 0 jobs claimed, 0 done; 7 of 7 workers parked, 7 sleeping; mode 0
expert pool threads: w0=sleeping ... w6=sleeping; host idle for 198437 ms
verify window (last window, not the current stage): 6 tokens at position 0, host at layer step 1;
the GPU rang 1; flags: served 0, plan (A) 0, copies (B) 0
memory: 36612 MiB resident, 37063 MiB committed, 14896 MiB RAM available; 10220761 page faults so far
wrote the thread stacks to C:\Strata-0.1.40.1\strata-stall-13584.dmp (attach it to the issue)
strata: released the verify window's GPU waits (#267): the GPU finished in 65 ms
```
The last line looks like the key: **the GPU work is done in ~65–68 ms once the waits are force-released**, so the
host was waiting for a signal/handshake that never arrives in the normal path. The engine then exits with
`0xC0000409` and the server returns `503`; the next request starts the engine again and it stalls identically.
3. Deterministic: reproduced on every start, always at `layer step 1`, with `6 tokens at position 0`.
## Ruled out
* **Not RAM**: 14.9 GB RAM free during the IQ2_XS stall (8.6–10.4 GB during the IQ3_XXS runs); page file grew
itself from 3.8 GB to 7.5 GB.
* **Not the model**: identical behaviour with IQ3_XXS and IQ2_XS; the *same* IQ2_XS shards run in llama.cpp
(build 11457) on the same PC/GPU: prompt 11–16 tok/s, decode ~10 tok/s, two back-to-back long generations
(783 + 398 tokens) with no stall.
* **Not the disk**: model and pack on NVMe (33 GiB loaded at 2.46 GiB/s); also reproduced with the model on
another volume.
* **Not the GPU being busy**: `nvidia-smi` shows 0% utilisation while the host waits; 474 MiB VRAM free with
everything loaded.
* **Not a driver reset**: no `nvlddmkm`/TDR events in the Windows event log (last 3 days).
## Configurations/engine flags tried (all stall the same way)
1. setup default config (64K context, `--kv int8`, `--expert-cache auto`, `--prefill auto`, `--spec 4`, MTP)
2. `--host-core last` (Windows sends GPU interrupts to one logical processor; default puts the host thread on 0)
3. `STRATA_HC_SPLIT=0` (the non-staged hyper-connection read)
4. `--short-read 0` — **changes where it stops**: the prompt is then processed via the batched path and the stall
moves to the first *decode* window (`1 tokens at position 14, host at layer step 1`), same signature
5. `--no-hit-poke`
6. `--expert-cache 0` (the engine still auto-sized a 5618-slot cache from the profile)
7. `STRATA_TOPK_STREAM=0` (the 0.1.40 Turing escape hatch from TROUBLESHOOTING.md)
8. `--pcie-mode dma` (the pre-0.1.14 expert-copy path from issue #31)
9. `--no-token-graph` → refused by `--serve`: `needs --spec T and --prefill CHUNK (and a fillable --expert-cache; the graphed hit path additionally needs --expert-profile P)`
10. `--spec 1` → refused: `--spec T (T >= 2)`
11. IQ2_XS instead of IQ3_XXS (this run also had `--kv-resident 32768`)
Also seen at startup in every run (probably unrelated, but mentioning it):
```
strata generate: expert arena: cudaHostRegister PORTABLE ok; large pages refused for 35456548864 B
(GetLargePageMinimum=2097152, VirtualAlloc error 1450); using 4 KB pages
```
[strata-stall-21656.dmp](https://github.com/user-attachments/files/33210397/strata-stall-21656.dmp)
## Attachments
* `strata-iq2_xs.log` (20 KB) — the IQ2_XS run, includes the stall report above
* `strata-iq3_xxs.log` (57 KB) — the IQ3_XXS run, same stall
* `strata-stall-*.dmp` — 13 files, 0.31–0.33 MB each (thread stacks; pick any one or two)
## Guess at the cause (for what it is worth)
`host at layer step 1` + `the GPU rang 1` + "the GPU finished in 65 ms after the waits were released" points at the
GPU handshake in the verify-window path, i.e. the host waits forever for a completion. Engine 0.1.14 fixed a
similar "host waits forever inside the NVIDIA driver" case for IQ packs (#31); this looks like another instance of
it on Turing + driver 616.92. Please tell me if a debug build, a `--sync-every-layer` run, or an older/newer
engine build would give you more useful data — I can rerun and send whatever you need.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.