Issues / #649

#649 HIP (gfx1030): intermittent 'verify: timed out at layer N (#267)' with UD-Q4_K_XL — repro data

open · @jollyroger1480 · 3 commentaires · Sur GitHub

AMD / HIPModels & quants

Description

## Environment

- RX 6950 XT (gfx1030), ROCm 7.1, HIP backend, engine **0.1.37** and **current main** (99f3dbd) — same behavior on both
- 31 GB RAM, UD-Q4_K_XL (`--native` GGUF-in-place, `--resident-budget-gib`), `--spec 2`, `STRATA_HC_SPLIT=0` (the default hc staging path also hangs on this card)
- Model: the Huihui abliterated re-split of Qwen3.8-Flash-Next UD-Q4_K_XL (needs the split-shard arch-guard fix from #648 to load at all)

## Symptom

Decode usually starts and can generate 24+ tokens correctly, then a verify window dies:

```
strata generate: verify: timed out at layer 11; its GPU waits were released but the GPU did not finish within 5 s (#267)
```

- Always at a full-attention layer (3, 11, 13, 25, 35, 46 seen — every N was (L+1)%4==0).
- **Intermittent**: same command, same budget: one run completes 24 tokens, the next hangs on the first verify round.
- When the release message says 'the GPU finished' the spin ended at release; when it says 'did not finish within 5 s' the graph was still running 25 s in.
- Frequency scales with `--resident-budget-gib`: 12 GiB hung 2/2 runs; 7 GiB roughly 1-in-3.

## Ruled out

- `--sync-every-layer` — still hangs
- `STRATA_ADAPT_NOWAIT=1` (main) — still hangs
- `--no-pool` — different failure (`--spec needs the device residency table`), so the pool interleave cannot be fully bypassed
- It is **not** the #648 loader issue — that is fixed in these runs (prefill + many decode rounds complete before the hang)

## Suspect

Timing-sensitive flag handshake between the CPU expert pool and the verify window's spin kernels — the warm-page-cache runs (fast PLE, fast prefill) hang more often than the first cold run, which completed. Happy to run any instrumented build or capture rocm-smi/rid traces if useful.

Sur le site

Liens install, modèles, releases.