反馈 / #1709

#1709 [Windows 10][AMD RX 9070 XT][HIP] Strata 0.1.41 loads IQ3_XXS but crashes on every generation request

open · @fifawot710-hub · 0 评论 · 去 GitHub 看

Setup & installServer & APIAMD / HIPModels & quantsWindows

说明

<html>
<body>
<!--StartFragment--><html><head></head><body><h2><span>Summary</span></h2><p><span>Strata 0.1.41 on Windows 10 with an AMD RX 9070 XT (gfx1201) successfully loads Qwen3.8-Flash-Next IQ3_XXS but fails on every generation request, including a minimal </span><code><span>Hi</span></code><span> request through the OpenAI-compatible API.</span></p><p><span>Depending on the environment settings, the engine crashes with </span><code><span>0xC0000005</span></code><span>, fails during prefill, or returns </span><code><span>verify: layer N never rang (graph finished ...)</span></code><span>.</span></p><p><span>No configuration tested so far has produced a successful response.</span></p><h2><span>System configuration</span></h2>
Component | Specification
-- | --
OS | Windows 10
CPU | AMD Ryzen 7 9800X3D
GPU | AMD Radeon RX 9070 XT, 16 GB (gfx1201)
RAM | 64 GB DDR5-6000
GPU driver | AMD Adrenalin 26.8.1
Strata | Engine 0.1.41, Windows HIP backend
Model | Qwen3.8-Flash-Next GSQ-RCO IQ3_XXS
Context | 65,536 tokens
KV cache | int8
Original speculative decoding | --spec 4 --spec-min-p 0.5

<h2><span>Steps to reproduce</span></h2><ol start="1"><li><p><span>Install Strata using </span><code><span>START-HERE.bat</span></code><span> and select Qwen3.8-Flash-Next IQ3_XXS with the HIP backend.</span></p></li><li><p><span>Wait until the model finishes loading and the server reports </span><code><span>ready</span></code><span>.</span></p></li><li><p><span>Open the web UI and send a simple prompt, such as </span><code><span>Hi</span></code><span>.</span></p></li><li><p><span>The engine terminates unexpectedly without producing a response.</span></p></li></ol><p><span>The same behavior can be reproduced using the API:</span></p><pre><code><span>{
  "model": "anything",
  "messages": [
    {"role": "user", "content": "Hi"}
  ],
  "max_tokens": 32,
  "stream": false
}</span></code></pre><p><strong><span>Expected:</span></strong><span> A normal generated response.</span></p><p><strong><span>Actual:</span></strong><span> The engine crashes, or the API returns an error during verification.</span></p><h2><span>Observed failures</span></h2><h3><span>1. Initial HIP access violation</span></h3><p><span>With the original configuration:</span></p><pre><code><span>prefill gemm: hipBLASLt tuning enabled (32 rows, gfx1201, version 100500)
strata serve: 296 MiB of VRAM free with everything loaded
Exception Code: 0xC0000005
amdhip64_7.dll ... hipModuleLoad()
libtensilelite-host.dll ... loadCodeObjectFile()
libhipblaslt.dll ... hipblasLtMatmul()</span></code></pre><p><span>Windows reports exit code </span><code><span>3221225477</span></code><span>.</span></p><h3><span>2. After increasing VRAM reserve and changing BLAS settings</span></h3><p><span>With </span><code><span>--vram-reserve-mib 2048</span></code><span> and </span><code><span>ROCBLAS_USE_HIPBLASLT=0</span></code><span>:</span></p><pre><code><span>strata serve: 1718 MiB of VRAM free with everything loaded
prefill mmq: mul_mat_q: unspecified launch failure</span></code></pre><p><span>The engine exits with code 1.</span></p><h3><span>3. After disabling MMQ prefill</span></h3><p><span>With </span><code><span>STRATA_PREFILL_MMQ=0</span></code><span>:</span></p><pre><code><span>prefill gemm: cublasGemmEx f16: cuBLAS status 6</span></code></pre><p><span>Another run produced:</span></p><pre><code><span>prefill gemm: cublasGemmEx f16 left unspecified launch failure</span></code></pre><h3><span>4. Direct API requests</span></h3><p><span>A minimal </span><code><span>Hi</span></code><span> request fails with:</span></p><pre><code><span>verify: layer 19 never rang (graph finished ...)</span></code></pre><p><span>After changing </span><code><span>--spec 4</span></code><span> to </span><code><span>--spec 2</span></code><span>, the error persists:</span></p><pre><code><span>verify: layer 20 never rang (graph finished ...)</span></code></pre><h2><span>Troubleshooting already performed</span></h2><ul><li><p><span>Moved installation and model directories to ASCII-only paths to resolve an earlier file-loading problem.</span></p></li><li><p><span>Increased VRAM reserve from 700 to 2048 MiB.</span></p></li><li><p><span>Disabled hipBLASLt tuning using </span><code><span>STRATA_HIPBLASLT_TUNING=""</span></code><span>.</span></p></li><li><p><span>Set </span><code><span>ROCBLAS_USE_HIPBLASLT=0</span></code><span>.</span></p></li><li><p><span>Disabled MMQ prefill using </span><code><span>STRATA_PREFILL_MMQ=0</span></code><span>.</span></p></li><li><p><span>Tested speculative decoding with </span><code><span>--spec 4</span></code><span> and </span><code><span>--spec 2</span></code><span>.</span></p></li><li><p><span>Tested both the web UI and a direct PowerShell API request.</span></p></li><li><p><span>Confirmed that the model loads successfully, but generation does not complete.</span></p></li></ul><p><span>All attempts failed. The specific error changes depending on the selected compute path.</span></p><h2><span>Related issues</span></h2><ul><li><p><span>#1644 — Windows HIP with IQ3_XXS on RX 7900 XT; model loads, but generation fails.</span></p></li><li><p><span>#613 — RX 9070 XT with Adrenalin 26.8.1; GPU timeout during generation.</span></p></li></ul><p><span>These reports have related symptoms but are not necessarily the same underlying bug.</span></p><h2><span>Additional information</span></h2><p><span>The installation uses the prebuilt Windows HIP engine with bundled ROCm libraries.</span></p><p><span>The model weights and expert cache initialize successfully. Approximately 39.97 GiB of experts are loaded into system RAM.</span></p><p><span>I can provide sanitized engine logs and run additional targeted diagnostic tests if needed.</span></p><p><span>Is this a known compatibility problem with the Windows HIP backend on gfx1201? Are there recommended diagnostic flags or a specific engine build that could help isolate the failing GPU operation?</span></p></body></html><!--EndFragment-->
</body>
</html>

本站相关内容

相关页面的快捷入口。