Issues / #613
#613 AMD 9070 XT on Windows causes VIDEO_ENGINE_TIMEOUT_DETECTED
open · @galmok · 7 comentários · No GitHub
BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentationWindows
Descrição
I just clones the repo, ran START-.bat, answered questions, and a browser window showed and I entered a few standard queries (tell me a joke, tell me a story) and I got a usable response at 54 tps but on the third query (tell me a longer story), the GPU drivers crashed, triggering a .dmp file that reports this error: VIDEO_ENGINE_TIMEOUT_DETECTED
Hardware:
AMD 98950X3D
128GB RAM
AMD 9070XT
Software:
Windows 11.
GPU driver: 26.8.1
This is my full command line history:
> PS C:\Tools\Strata\Strata> .\START-HERE.bat
> Strata - Qwen3.8-Flash-Next on a normal PC (a GPU + system RAM + CPU)
>
> === Step 1: checking your PC ===
> Your AMD GPUs:
> GPU 0: AMD Radeon RX 9070 XT, 16 GB VRAM - can be used
> GPU 1: AMD Radeon(TM) Graphics, 2 GB VRAM - not supported - Strata's AMD backend runs on the RX 7900 XT / XTX (gfx1100), RX 7800 XT / 7700 XT (gfx1101), RX 9060 XT (gfx1200) and RX 9070 / 9070 XT / Radeon AI PRO R9700 (gfx1201), and the RX 6800 / 6900 series (gfx1030, unvalidated) only, this is unknown (PCI 13C0)
> [ok] GPU: AMD Radeon RX 9070 XT, 15.9 GB VRAM, gfx1201 (AMD: docs/AMD_HIP.md)
> [ok] RAM: 126 GB
> [ok] CPU: AMD64 Family 26 Model 68 Stepping 0, AuthenticAMD (AVX-512)
>
> === Step 2: your choices ===
> 1) Qwen3.8-Flash-Next Qwen; GSQ-RCO quants by ISTA-DASLab - the original model
> 2) Swift 1.5 UkisAI's fine-tune of Qwen3.8-Flash-Next - thinks much shorter (-63% thinking tokens, 1.8x sooner answers by its authors' numbers)
> 3) Qwen3.8-Flash-Next Coder ISTA-DASLab's coding version - half the experts (code, tools, images kept): needs ~32 GB of RAM, faster; weaker outside coding
> 4) Qwen3.8-Flash-Next (Unsloth) Unsloth's 4-bit quantization (EXPERIMENTAL) - 4-bit, 111 GB download, most experts read from the SSD: slow (7-8.5 tokens/s on a 64 GB PC) [experimental]
> Which model? [1]: 1
> [ok] model: Qwen3.8-Flash-Next
>
> 1) Q2_0 2-bit, the fastest; download 66 GB, uses ~34 GB of RAM
> 2) IQ2_XS 2-bit i-quant, a little better quality, close in speed; download 68 GB, uses ~36 GB of RAM
> 3) IQ3_XXS 3-bit i-quant, better quality, slower (more CPU work per token); download 76 GB, uses ~43 GB of RAM
> 4) IQ3_S 3.5-bit i-quant, the best quality (matches the full model), the slowest; needs a 64 GB PC with little else running; download 84 GB, uses ~50 GB of RAM
> Which size? [3]: 4
> [ok] size: IQ3_S
>
> Context length = how much text the model can see at once (your chat, files, tool output).
> Longer needs more VRAM for it, so fewer experts fit on the GPU:
> 1) 8K tokens
> 2) 32K tokens
> 3) 64K tokens (recommended for your GPU)
> 4) 128K tokens
> 5) 256K tokens
> 6) 384K tokens (experimental: setup adds rope scaling)
> 7) 512K tokens (experimental: setup adds rope scaling)
> Context? [3]: 5
> [ok] context: 262144 tokens
>
> KV cache precision (the model's memory of the conversation):
> 1) 8-bit (recommended: what every published number was measured with)
> 2) 4-bit half the memory (about 4% faster at 128K), but measurably less precise on long
> documents; long-context lookups (needle tests) still pass
> KV cache? [1]: 1
> [ok] KV cache: 8-bit
> [ok] images: off
>
> EXPERIMENTAL - speed projection: a small control vector applied while the model runs (layers 4-44).
> It changes how the model answers: its package describes it as a refusal-direction projection (the
> model declines far fewer requests). Off unless you choose it; when on, the web app can switch it off
> per chat. Details: data/experimental-speed-projection/README.md
> Turn on the experimental speed projection? [n]:
> [ok] experimental speed projection: off
>
> === Step 3: Python packages ===
> Installing numpy, jinja2, regex, pyyaml, tqdm, requests, cmake, ninja, pillow, psutil ...
> > C:\Tools\Strata\Strata\.venv\Scripts\python.exe -m pip install --quiet --disable-pip-version-check numpy==2.5.3; python_version >= "3.12" numpy==2.4.6; python_version == "3.11" numpy==2.2.6; python_version < "3.11" jinja2==3.1.6 regex==2026.9.10 pyyaml==6.0.3 tqdm==4.70.1 requests==2.34.2 cmake==4.4.3 ninja==1.13.2 pillow==12.3.0 psutil==7.2.2 markupsafe==3.0.3 certifi==2026.7.22 charset-normalizer==3.5.1 idna==3.20 urllib3==2.8.0 colorama==0.4.6; sys_platform == "win32"
> [ok] numpy, jinja2, regex, pyyaml, tqdm, requests, cmake, ninja, pillow, psutil installed
>
> === Step 4: the Strata engine ===
> llama.cpp source: 8.4 MB
> [ok] llama.cpp source downloaded
> [ok] llama.cpp 3cf0325 (gguf-py, ggml, mtmd)
> Downloading the ready-made Strata engine for AMD GPUs (with the ROCm libraries it uses) ...
> Strata AMD engine: 0.49 / 0.60 GB (83%)
> [ok] Strata AMD engine downloaded
> [ok] ready-made AMD engine 0.1.38 for gfx1100, gfx1101, gfx1102, gfx1200, gfx1201, gfx1030 (ROCm 10.2.0a20260930)
> [ok] HIP numbers this card 1 (an integrated GPU comes first): the engine is pointed at it
> [ok] engine: C:\Tools\Strata\Strata\engine\strata.exe
>
> === Step 5: downloading Qwen3.8-Flash-Next IQ3_S ===
> The model files go in C:\Tools\Strata\Strata-data\models\IQ3_S
> Files you already have: put them here with their original names (Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf, Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf), or use --gguf-dir <their folder>.
> Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf: 54.75 / 54.82 GB (100%)
> [ok] Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf downloaded
> Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf: 28.80 / 28.80 GB (100%)
> [ok] Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf downloaded
> [ok] model files present
>
> === Step 6: preparing the model for Strata ===
> > C:\Tools\Strata\Strata\.venv\Scripts\python.exe C:\Tools\Strata\Strata\tools\iq_pack.py --gguf C:\Tools\Strata\Strata-data\models\IQ3_S\Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf --out C:\Tools\Strata\Strata-data\packs\iq3_s
> model shards: Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf, Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf
> index.txt: 1079 tensors, 302 served natively, 0 in extra.bin, arena 1.43 GiB
> tokenizer/: vocab 248320, merges 247587, pre qwen35, specials {'tokenizer.ggml.eos_token_id': 248046, 'tokenizer.ggml.padding_token_id': 248044, 'tokenizer.ggml.bos_token_id': 248044}
> [ok] model prepared: C:\Tools\Strata\Strata-data\packs\iq3_s
> The MTP draft layer (speculative decoding, ~2x faster output) comes from the original Qwen checkpoint:
> only its ~5 GB of MTP tensors are downloaded.
> > C:\Tools\Strata\Strata\.venv\Scripts\python.exe C:\Tools\Strata\Strata\tools\mtp_fetch.py fetch --out C:\Tools\Strata\Strata-data\mtp
> # MTP block in the BF16 checkpoint
>
> 31 tensors, 5.214 GB, in 28 shards.
>
> | tensor | dtype | shape | MB |
> |---|---|---|---:|
> | `mtp.fc_embedding.weight` | BF16 | 2560x2560 | 13.1 |
> | `mtp.fc_hidden.weight` | BF16 | 2560x2560 | 13.1 |
> | `mtp.hyper_connection_mixer.hc_norm.weight` | BF16 | 10240 | 0.0 |
> | `mtp.hyper_connection_mixer.input_mix_weight_down.weight` | BF16 | 320x10240 | 6.6 |
> | `mtp.hyper_connection_mixer.input_mix_weight_up.weight` | BF16 | 10240x320 | 6.6 |
> | `mtp.layers.0.attn_hyper_connection.block_inject_weight.weight` | BF16 | 4x10240 | 0.1 |
> | `mtp.layers.0.attn_hyper_connection.hc_norm.weight` | BF16 | 10240 | 0.0 |
> | `mtp.layers.0.attn_hyper_connection.input_mix_weight_down.weight` | BF16 | 320x10240 | 6.6 |
> | `mtp.layers.0.attn_hyper_connection.input_mix_weight_up.weight` | BF16 | 10240x320 | 6.6 |
> | `mtp.layers.0.mlp.experts.down_proj` | BF16 | 512x2560x640 | 1677.7 |
> | `mtp.layers.0.mlp.experts.gate_up_proj` | BF16 | 512x1280x2560 | 3355.4 |
> | `mtp.layers.0.mlp.gate.weight` | BF16 | 512x2560 | 2.6 |
> | `mtp.layers.0.mlp.shared_expert.down_proj.weight` | BF16 | 2560x640 | 3.3 |
> | `mtp.layers.0.mlp.shared_expert.gate_proj.weight` | BF16 | 640x2560 | 3.3 |
> | `mtp.layers.0.mlp.shared_expert.up_proj.weight` | BF16 | 640x2560 | 3.3 |
> | `mtp.layers.0.mlp.shared_expert_gate.weight` | BF16 | 1x2560 | 0.0 |
> | `mtp.layers.0.mlp_hyper_connection.block_inject_weight.weight` | BF16 | 4x10240 | 0.1 |
> | `mtp.layers.0.mlp_hyper_connection.hc_norm.weight` | BF16 | 10240 | 0.0 |
> | `mtp.layers.0.mlp_hyper_connection.input_mix_weight_down.weight` | BF16 | 320x10240 | 6.6 |
> | `mtp.layers.0.mlp_hyper_connection.input_mix_weight_up.weight` | BF16 | 10240x320 | 6.6 |
> | `mtp.layers.0.self_attn.indexer.index_qk_proj.weight` | BF16 | 640x2560 | 3.3 |
> | `mtp.layers.0.self_attn.indexer.k_layernorm.weight` | BF16 | 128 | 0.0 |
> | `mtp.layers.0.self_attn.indexer.q_layernorm.weight` | BF16 | 128 | 0.0 |
> | `mtp.layers.0.self_attn.k_norm.weight` | BF16 | 256 | 0.0 |
> | `mtp.layers.0.self_attn.k_proj.weight` | BF16 | 512x2560 | 2.6 |
> | `mtp.layers.0.self_attn.o_proj.weight` | BF16 | 2560x6144 | 31.5 |
> | `mtp.layers.0.self_attn.q_norm.weight` | BF16 | 256 | 0.0 |
> | `mtp.layers.0.self_attn.q_proj.weight` | BF16 | 12288x2560 | 62.9 |
> | `mtp.layers.0.self_attn.v_proj.weight` | BF16 | 512x2560 | 2.6 |
> | `mtp.pre_fc_norm_embedding.weight` | BF16 | 2560 | 0.0 |
> | `mtp.pre_fc_norm_hidden.weight` | BF16 | 10240 | 0.0 |
>
> mtp.fc_embedding.weight 100%
> mtp.fc_hidden.weight 100%
> mtp.hyper_connection_mixer.hc_norm.weight 100%
> mtp.hyper_connection_mixer.input_mix_weight_down.weight 100%
> mtp.hyper_connection_mixer.input_mix_weight_up.weight 100%
> mtp.layers.0.attn_hyper_connection.block_inject_weight.weight 100%
> mtp.layers.0.attn_hyper_connection.hc_norm.weight 100%
> mtp.layers.0.attn_hyper_connection.input_mix_weight_down.weight 100%
> mtp.layers.0.attn_hyper_connection.input_mix_weight_up.weight 100%
> mtp.layers.0.mlp.experts.down_proj 4%
> mtp.layers.0.mlp.experts.down_proj 8%
> mtp.layers.0.mlp.experts.down_proj 12%
> mtp.layers.0.mlp.experts.down_proj 16%
> mtp.layers.0.mlp.experts.down_proj 20%
> mtp.layers.0.mlp.experts.down_proj 24%
> mtp.layers.0.mlp.experts.down_proj 28%
> mtp.layers.0.mlp.experts.down_proj 32%
> mtp.layers.0.mlp.experts.down_proj 36%
> mtp.layers.0.mlp.experts.down_proj 40%
> mtp.layers.0.mlp.experts.down_proj 44%
> mtp.layers.0.mlp.experts.down_proj 48%
> mtp.layers.0.mlp.experts.down_proj 52%
> mtp.layers.0.mlp.experts.down_proj 56%
> mtp.layers.0.mlp.experts.down_proj 60%
> mtp.layers.0.mlp.experts.down_proj 64%
> mtp.layers.0.mlp.experts.down_proj 68%
> mtp.layers.0.mlp.experts.down_proj 72%
> mtp.layers.0.mlp.experts.down_proj 76%
> mtp.layers.0.mlp.experts.down_proj 80%
> mtp.layers.0.mlp.experts.down_proj 84%
> mtp.layers.0.mlp.experts.down_proj 88%
> mtp.layers.0.mlp.experts.down_proj 92%
> mtp.layers.0.mlp.experts.down_proj 96%
> mtp.layers.0.mlp.experts.down_proj 100%
> mtp.layers.0.mlp.experts.gate_up_proj 2%
> mtp.layers.0.mlp.experts.gate_up_proj 4%
> mtp.layers.0.mlp.experts.gate_up_proj 6%
> mtp.layers.0.mlp.experts.gate_up_proj 8%
> mtp.layers.0.mlp.experts.gate_up_proj 10%
> mtp.layers.0.mlp.experts.gate_up_proj 12%
> mtp.layers.0.mlp.experts.gate_up_proj 14%
> mtp.layers.0.mlp.experts.gate_up_proj 16%
> mtp.layers.0.mlp.experts.gate_up_proj 18%
> mtp.layers.0.mlp.experts.gate_up_proj 20%
> mtp.layers.0.mlp.experts.gate_up_proj 22%
> mtp.layers.0.mlp.experts.gate_up_proj 24%
> mtp.layers.0.mlp.experts.gate_up_proj 26%
> mtp.layers.0.mlp.experts.gate_up_proj 28%
> mtp.layers.0.mlp.experts.gate_up_proj 30%
> mtp.layers.0.mlp.experts.gate_up_proj 32%
> mtp.layers.0.mlp.experts.gate_up_proj 34%
> mtp.layers.0.mlp.experts.gate_up_proj 36%
> mtp.layers.0.mlp.experts.gate_up_proj 38%
> mtp.layers.0.mlp.experts.gate_up_proj 40%
> mtp.layers.0.mlp.experts.gate_up_proj 42%
> mtp.layers.0.mlp.experts.gate_up_proj 44%
> mtp.layers.0.mlp.experts.gate_up_proj 46%
> mtp.layers.0.mlp.experts.gate_up_proj 48%
> mtp.layers.0.mlp.experts.gate_up_proj 50%
> mtp.layers.0.mlp.experts.gate_up_proj 52%
> mtp.layers.0.mlp.experts.gate_up_proj 54%
> mtp.layers.0.mlp.experts.gate_No site
Links install, modelos, releases.