Issues / #1691

#1691 gfx906: 0.1.41 verify timeouts ("timed out at layer N"); prompt-read stalls are pre-existing (clamp blame retracted)

open · @KevinX8 · 1 commentaires · Sur GitHub

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

Description

Bisected a gfx906 regression to the q8_1 clamp pair. Reporter setup below, A/B table, then logs.

## Setup

- 1x AMD Instinct MI50 16 GB (gfx906, wave64), Ryzen 9 5900X (AVX2, no AVX-512), 60 GB RAM, Ubuntu 26.10, kernel amdgpu
- ROCm 10.1 community gfx906 build (also reproduced on 10.0), HIP 7.16, clang 24
- Engine built with `-DSTRATA_HIP_GFX906=ON -DCMAKE_HIP_ARCHITECTURES=gfx906` (docs/AMD_HIP.md recipe)
- Model: Unsloth UD-IQ4_XS (native pack), 262144 context, q4_0 KV + `--kv-resident 32768`, `--resident-budget-gib 43`, MTP `--spec 4 --spec-min-p 0.5`, single GPU
- 0.1.40.x needs one local patch to build (two defines in the gfx906 compat header, same class as #1396):
  `cudaStreamCreateWithPriority` / `cudaDeviceGetStreamPriorityRange` in `include/strata/platform/hip_compat/strata_hip.h` (a732a19 fixed the neighbouring `cudaEventBlockingSync` but not these; `src/core/mtp.cpp:423-424` uses them)

## What 0.1.41 does (fb58e0d)

Two failure signatures on the same box where 0.1.40.x serves all day:
1. `Unrecoverable native verification failure: verify: timed out at layer 10` and `... layer 26; its GPU waits were released` — decode stalls (18 -> 4.7 tok/s) then the engine exits, twice.
2. `no progress for 60 s during a request (reading the prompt (batched): waiting for the GPU (attention, router) ...` — twice more.

Decode on the runs that complete: ~19.8 tok/s vs ~26-28 tok/s on 0.1.40.x.

## Bisect (v0.1.40.4 base, same flags/config/machine)

- v0.1.40.4 alone: story x3 + a full live session, no hang, 23-28 tok/s decode.
- v0.1.40.4 + ccd657a (HC norm scratch): 3 stories, 29.0/27.6 tok/s after warmup, no hang. Exonerated.
- v0.1.40.4 + e372d4f + 7a0e302 (q8_1 clamps, isolated, no ccd657a): 4 consecutive prompt-read hangs across mixed traffic — a 27-token prompt dying 12 tokens in, a 2.6K prompt dying 5 tokens from the end, a 3.8K prompt dying 5 tokens from the end — every one `waiting for the GPU (attention, router)`, each ended by the 60 s watchdog + engine restart. Nothing else served in between except the hangs.

So the clamp pair alone reproduces a hang on gfx906. The 0.1.41 verify-timeout crashes were not re-tested to 5x (stopped per the reporter's rule after the bisect landed); they may be the same root cause surfacing in decode.

Happy to re-run anything on this card (single MI50 16 GB, can build any commit with the one-line patch above).

Sur le site

Liens install, modèles, releases.