Issues / #954
#954 [BUG] 0.1.39: fused int8 prompt experts (STRATA_PF_FUSED=1) + #583 byte-budget ring → kernel hang, `prefill mmq: iota: unknown error` on sm_120 (RTX 5090, WSL2)
open · @Krypto-Whitehat · 3 comentarios · En GitHub
BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsLinux
Descripción
## Summary
On engine 0.1.39, `STRATA_PF_FUSED=1` on a native IQ3_XXS pack hangs a CUDA kernel (GPU pegged at 100%, host blocked) during the first prompt, then the engine dies with `prefill mmq: iota: unknown error` (exit code 1). The same configuration ran ~4,700 successful requests on 0.1.37. Disabling the #583 byte-budget ring (`STRATA_RING_BYTES=0`) while keeping `STRATA_PF_FUSED=1` makes 0.1.39 work again, so the regression is in the ring's interaction with the fused prompt path, not the fused kernels themselves.
## Environment
- Engine: 0.1.39 (`6f32ec0`, built 2026-10-04, CUDA 13.0 build)
- GPU: NVIDIA GeForce RTX 5090 Laptop, compute capability 12.0 (Blackwell), 24 GiB VRAM
- Driver: KMD 617.14 / CUDA UMD 13.4, WSL2 (Ubuntu 24.04, kernel 6.18.33.2-microsoft-standard-WSL2), WSL 2.7.11.0
- Model: Qwen3.8-Flash-Next IQ3_XXS native pack (125B MoE), expert arena 39.97 GiB resident in hugetlb RAM, expert cache auto → 6116 slots (13.27 GiB VRAM), `--prefill auto:16384` → `prompt chunk auto: 16384 tokens, a 512-slot ring`, `--kv int8`, `--mtp` spec 4
- 277 MiB VRAM free after full load
## Reproduction
Config env `STRATA_PF_FUSED=1` (default ring), serve or generate mode, any prompt:
```
strata generate: prompt chunk auto: 16384 tokens, a 512-slot ring
strata serve: the prompt path borrows 2980 CUDA0 cache slots (6.47 GiB)
... engine dies: "prefill mmq: iota: unknown error" (exit code 1)
```
Observed as a hang first: with `CUDA_LAUNCH_BLOCKING=1` the host blocks inside a CUDA call while the GPU sits at 100% utilization indefinitely (14+ min) — consistent with a stuck kernel rather than a clean fault; `iota` is just the first `ck(cudaGetLastError())` checkpoint that surfaces the sticky error.
Control runs on the same machine/binary:
| env | result |
|---|---|
| `STRATA_PF_FUSED=1` (ring default on) | hang / `prefill mmq: iota: unknown error` |
| `STRATA_PF_FUSED=0` (ring on) | works — 7260-token prompt at ~1397 tok/s prefill |
| `STRATA_PF_FUSED=1` + `STRATA_RING_BYTES=0` | works — fused marker logs, prompt + decode complete |
0.1.37 history (same log file): `session is up (engine 0.1.37)` + `prompt experts on the fused int8 kernels` markers followed by thousands of completed requests — fused worked fine before the f-139b merge set.
## Suspected area
#583 commits between the builds: `895a77b` (streamed ring is a byte budget), `a93c2ac` (chunk and ring share one budget), `8c968a5`/`5423a69`/`5f19911` (ring on by default). Possibly also `220e0e8` (fused layout buffer sizing gating on `fused_ring()`). Whatever changed the ring's sizing/lifetime now gives the fused kernels a bad buffer — worth checking ring slot ownership when `fused_layout()` shrinks GU/H/Xq while the byte-budget ring still assumes full-size rows.
## Workaround
`"env": {"STRATA_PF_FUSED": "1", "STRATA_RING_BYTES": "0"}` in `strata-*.json` — fused prompt experts stay on and work, the byte-budget ring is off.
## Impact
High for fused-path users on Blackwell: deterministic engine death on first prompt (and the WDDM driver then reports `GPU access blocked by the operating system` to nvidia-smi until the WSL VM is restarted, masking the failure as a GPU-detection problem).
En el sitio
Enlaces a install, modelos, releases.