Issues / #1233
#1233 ``` Windows + Strix Halo (gfx1151, Radeon 8060S): UD-IQ4_XS runs end-to-end — decode flat 29.5-30.4 tok/s at 1K-32K, but a GPU driver timeout (TDR) during a cold 32K prefill ```
open · @Facostarr · 4 Kommentare · Auf GitHub
BenchmarksSetup & installAMD / HIPModels & quantsWindowsLinux
Beschreibung
```markdown **First end-to-end real-model run on Windows + gfx1151.** Issue #918 covers building, self-tests and the HIP ctest on this chip; as far as I can find, this is the first report of an actual model served on a Radeon 8060S under Windows. Related: #612 (feature request), #613 (a TDR on Windows AMD, discrete card), #953 (the Windows/AMD report format), #1141 (the prebuilt HIP zip on a discrete card), #979 and #505 (hipBLASLt tables on AMD). ## Hardware / OS / build - Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB unified, 32 GB BIOS carve-out (Windows sees 95.6 GB). - Windows 11 26H2, build 26300.9457. GPU driver 32.0.31041.1004 (Adrenalin 26.8.1). Driver was NOT touched before or after this test. - Strata v0.1.40, ready-made `strata-windows-x64-hip.zip` (sha256 matched the published digest). `BUILD.json`: arch gfx1151, ROCm 10.2.0a20260930. - Setup choices: Unsloth **UD-IQ4_XS**, context **32K**, KV **int8**. Written settings: `--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --max-context 32768 --kv int8 --resident-budget-gib 55`, server on 127.0.0.1:8080. Strata reports the GPU as 104 GB. ## Measured (server-side timings, never wall clock; medians of 3 runs) | Context | Prompt tok/s | Output tok/s | | --- | --- | --- | | 1K | 104.2 | 29.5 | | 8K | 76.6 | 30.4 | | 32K | 72.1 (1 cold run) | 29.5 | The 32K prefill has a single cold sample — the 3-run pass was interrupted, see below. MTP draft on a warmup call: 13 drafted, 10 accepted. Two observations that may be useful to you: 1. **Decode is flat from 1K to 32K (29.5–30.4 tok/s).** The expert streaming path holds up well here. 2. **Prefill is 13–18x below your published Linux numbers on the same machine class** (1,293 / 1,370 / 1,320 tok/s on Ryzen AI Max+ 395, 128 GB, UD-IQ4_XS). ## GPU driver timeout (TDR) during a cold 32K prefill Window: 2026-10-06 ~19:38 local (UTC+2), during the first cold 32K prefill of the 3-run pass. AMD's tool raised: *"The AMD software detected a driver timeout on your system. A problem report has been created."* All earlier runs (decode 1K/8K/32K, prefill 1K/8K) completed normally. **The server survived**: `/health` kept answering `status: ok, loaded: true` afterwards. Windows System and Application logs showed **no 4101/4102/4103** (and no amd/display provider entry) over the following 60 minutes at Error/Warning/Critical level, so I cannot give you a classic TDR event id — only the AMD tool's own report. **Hypothesis, not a conclusion.** Setup printed: ``` [!] no hipBLASLt tuning table for gfx1151 with hipBLASLt 100500 (have: gfx1151-hipblaslt-100401.txt): the prompt's dense matrix products use plain hipBLAS ``` So the prefill's dense GEMMs take the unoptimised path while the context is at its largest and the expert cache is under most pressure. That would explain the prefill gap, and a long unoptimised kernel is a plausible candidate for the watchdog. I have **not** regenerated a table, and I have not tried `STRATA_HIP_PROMPT_F16=1` — tell me if either is worth a run and I will report back on the same hardware. For context on the table question: #979 reports a gfx1201 table for 100201 making prompts *slower* than no table, while #505 reports a 100202 table helping a lot on gfx1100 — so a gfx1151/100500 table is not obviously a win, which is why I did not spend the time unprompted. ## Caveats - Single machine, single session, no tuning of the engine's 18 auto-on gfx1151 switches. - A second engine (llama.cpp) baseline at the same context on the same box is being measured separately; I'll append it when it lands. - Happy to run any instrumented build or flag set you want tested on this exact configuration. ```
Mehr auf der Site
Links zu Install, Modellen, Releases.