Issues / #613

#613 AMD 9070 XT on Windows causes VIDEO_ENGINE_TIMEOUT_DETECTED

open · @galmok · 7 comments · View on GitHub

BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentationWindows

Description

I just clones the repo, ran START-.bat, answered questions, and a browser window showed and I entered a few standard queries (tell me a joke, tell me a story) and I got a usable response at 54 tps but on the third query (tell me a longer story), the GPU drivers crashed, triggering a .dmp file that reports this error: VIDEO_ENGINE_TIMEOUT_DETECTED

Hardware:

AMD 98950X3D
128GB RAM
AMD 9070XT

Software:

Windows 11.
GPU driver: 26.8.1

This is my full command line history:

> PS C:\Tools\Strata\Strata> .\START-HERE.bat
> Strata - Qwen3.8-Flash-Next on a normal PC (a GPU + system RAM + CPU)
> 
> === Step 1: checking your PC ===
>   Your AMD GPUs:
>     GPU 0: AMD Radeon RX 9070 XT, 16 GB VRAM - can be used
>     GPU 1: AMD Radeon(TM) Graphics, 2 GB VRAM - not supported - Strata's AMD backend runs on the RX 7900 XT / XTX (gfx1100), RX 7800 XT / 7700 XT (gfx1101), RX 9060 XT (gfx1200) and RX 9070 / 9070 XT / Radeon AI PRO R9700 (gfx1201), and the RX 6800 / 6900 series (gfx1030, unvalidated) only, this is unknown (PCI 13C0)
>   [ok] GPU: AMD Radeon RX 9070 XT, 15.9 GB VRAM, gfx1201 (AMD: docs/AMD_HIP.md)
>   [ok] RAM: 126 GB
>   [ok] CPU: AMD64 Family 26 Model 68 Stepping 0, AuthenticAMD (AVX-512)
> 
> === Step 2: your choices ===
>   1) Qwen3.8-Flash-Next   Qwen; GSQ-RCO quants by ISTA-DASLab - the original model
>   2) Swift 1.5            UkisAI's fine-tune of Qwen3.8-Flash-Next - thinks much shorter (-63% thinking tokens, 1.8x sooner answers by its authors' numbers)
>   3) Qwen3.8-Flash-Next Coder ISTA-DASLab's coding version - half the experts (code, tools, images kept): needs ~32 GB of RAM, faster; weaker outside coding
>   4) Qwen3.8-Flash-Next (Unsloth) Unsloth's 4-bit quantization (EXPERIMENTAL) - 4-bit, 111 GB download, most experts read from the SSD: slow (7-8.5 tokens/s on a 64 GB PC)   [experimental]
> Which model? [1]: 1
>   [ok] model: Qwen3.8-Flash-Next
> 
>   1) Q2_0     2-bit, the fastest; download 66 GB, uses ~34 GB of RAM
>   2) IQ2_XS   2-bit i-quant, a little better quality, close in speed; download 68 GB, uses ~36 GB of RAM
>   3) IQ3_XXS  3-bit i-quant, better quality, slower (more CPU work per token); download 76 GB, uses ~43 GB of RAM
>   4) IQ3_S    3.5-bit i-quant, the best quality (matches the full model), the slowest; needs a 64 GB PC with little else running; download 84 GB, uses ~50 GB of RAM
> Which size? [3]: 4
>   [ok] size: IQ3_S
> 
>   Context length = how much text the model can see at once (your chat, files, tool output).
>   Longer needs more VRAM for it, so fewer experts fit on the GPU:
>   1) 8K tokens
>   2) 32K tokens
>   3) 64K tokens   (recommended for your GPU)
>   4) 128K tokens
>   5) 256K tokens
>   6) 384K tokens   (experimental: setup adds rope scaling)
>   7) 512K tokens   (experimental: setup adds rope scaling)
> Context? [3]: 5
>   [ok] context: 262144 tokens
> 
>   KV cache precision (the model's memory of the conversation):
>   1) 8-bit   (recommended: what every published number was measured with)
>   2) 4-bit   half the memory (about 4% faster at 128K), but measurably less precise on long
>              documents; long-context lookups (needle tests) still pass
> KV cache? [1]: 1
>   [ok] KV cache: 8-bit
>   [ok] images: off
> 
>   EXPERIMENTAL - speed projection: a small control vector applied while the model runs (layers 4-44).
>   It changes how the model answers: its package describes it as a refusal-direction projection (the
>   model declines far fewer requests). Off unless you choose it; when on, the web app can switch it off
>   per chat. Details: data/experimental-speed-projection/README.md
> Turn on the experimental speed projection? [n]:
>   [ok] experimental speed projection: off
> 
> === Step 3: Python packages ===
>   Installing numpy, jinja2, regex, pyyaml, tqdm, requests, cmake, ninja, pillow, psutil ...
>   > C:\Tools\Strata\Strata\.venv\Scripts\python.exe -m pip install --quiet --disable-pip-version-check numpy==2.5.3; python_version >= "3.12" numpy==2.4.6; python_version == "3.11" numpy==2.2.6; python_version < "3.11" jinja2==3.1.6 regex==2026.9.10 pyyaml==6.0.3 tqdm==4.70.1 requests==2.34.2 cmake==4.4.3 ninja==1.13.2 pillow==12.3.0 psutil==7.2.2 markupsafe==3.0.3 certifi==2026.7.22 charset-normalizer==3.5.1 idna==3.20 urllib3==2.8.0 colorama==0.4.6; sys_platform == "win32"
>   [ok] numpy, jinja2, regex, pyyaml, tqdm, requests, cmake, ninja, pillow, psutil installed
> 
> === Step 4: the Strata engine ===
>   llama.cpp source:     8.4 MB
>   [ok] llama.cpp source downloaded
>   [ok] llama.cpp 3cf0325 (gguf-py, ggml, mtmd)
>   Downloading the ready-made Strata engine for AMD GPUs (with the ROCm libraries it uses) ...
>   Strata AMD engine:   0.49 / 0.60 GB (83%)
>   [ok] Strata AMD engine downloaded
>   [ok] ready-made AMD engine 0.1.38 for gfx1100, gfx1101, gfx1102, gfx1200, gfx1201, gfx1030 (ROCm 10.2.0a20260930)
>   [ok] HIP numbers this card 1 (an integrated GPU comes first): the engine is pointed at it
>   [ok] engine: C:\Tools\Strata\Strata\engine\strata.exe
> 
> === Step 5: downloading Qwen3.8-Flash-Next IQ3_S ===
>   The model files go in C:\Tools\Strata\Strata-data\models\IQ3_S
>   Files you already have: put them here with their original names (Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf, Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf), or use --gguf-dir <their folder>.
>   Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf:  54.75 / 54.82 GB (100%)
>   [ok] Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf downloaded
>   Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf:  28.80 / 28.80 GB (100%)
>   [ok] Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf downloaded
>   [ok] model files present
> 
> === Step 6: preparing the model for Strata ===
>   > C:\Tools\Strata\Strata\.venv\Scripts\python.exe C:\Tools\Strata\Strata\tools\iq_pack.py --gguf C:\Tools\Strata\Strata-data\models\IQ3_S\Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf --out C:\Tools\Strata\Strata-data\packs\iq3_s
> model shards: Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf, Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf
> index.txt: 1079 tensors, 302 served natively, 0 in extra.bin, arena 1.43 GiB
> tokenizer/: vocab 248320, merges 247587, pre qwen35, specials {'tokenizer.ggml.eos_token_id': 248046, 'tokenizer.ggml.padding_token_id': 248044, 'tokenizer.ggml.bos_token_id': 248044}
>   [ok] model prepared: C:\Tools\Strata\Strata-data\packs\iq3_s
>   The MTP draft layer (speculative decoding, ~2x faster output) comes from the original Qwen checkpoint:
>   only its ~5 GB of MTP tensors are downloaded.
>   > C:\Tools\Strata\Strata\.venv\Scripts\python.exe C:\Tools\Strata\Strata\tools\mtp_fetch.py fetch --out C:\Tools\Strata\Strata-data\mtp
> # MTP block in the BF16 checkpoint
> 
> 31 tensors, 5.214 GB, in 28 shards.
> 
> | tensor | dtype | shape | MB |
> |---|---|---|---:|
> | `mtp.fc_embedding.weight` | BF16 | 2560x2560 | 13.1 |
> | `mtp.fc_hidden.weight` | BF16 | 2560x2560 | 13.1 |
> | `mtp.hyper_connection_mixer.hc_norm.weight` | BF16 | 10240 | 0.0 |
> | `mtp.hyper_connection_mixer.input_mix_weight_down.weight` | BF16 | 320x10240 | 6.6 |
> | `mtp.hyper_connection_mixer.input_mix_weight_up.weight` | BF16 | 10240x320 | 6.6 |
> | `mtp.layers.0.attn_hyper_connection.block_inject_weight.weight` | BF16 | 4x10240 | 0.1 |
> | `mtp.layers.0.attn_hyper_connection.hc_norm.weight` | BF16 | 10240 | 0.0 |
> | `mtp.layers.0.attn_hyper_connection.input_mix_weight_down.weight` | BF16 | 320x10240 | 6.6 |
> | `mtp.layers.0.attn_hyper_connection.input_mix_weight_up.weight` | BF16 | 10240x320 | 6.6 |
> | `mtp.layers.0.mlp.experts.down_proj` | BF16 | 512x2560x640 | 1677.7 |
> | `mtp.layers.0.mlp.experts.gate_up_proj` | BF16 | 512x1280x2560 | 3355.4 |
> | `mtp.layers.0.mlp.gate.weight` | BF16 | 512x2560 | 2.6 |
> | `mtp.layers.0.mlp.shared_expert.down_proj.weight` | BF16 | 2560x640 | 3.3 |
> | `mtp.layers.0.mlp.shared_expert.gate_proj.weight` | BF16 | 640x2560 | 3.3 |
> | `mtp.layers.0.mlp.shared_expert.up_proj.weight` | BF16 | 640x2560 | 3.3 |
> | `mtp.layers.0.mlp.shared_expert_gate.weight` | BF16 | 1x2560 | 0.0 |
> | `mtp.layers.0.mlp_hyper_connection.block_inject_weight.weight` | BF16 | 4x10240 | 0.1 |
> | `mtp.layers.0.mlp_hyper_connection.hc_norm.weight` | BF16 | 10240 | 0.0 |
> | `mtp.layers.0.mlp_hyper_connection.input_mix_weight_down.weight` | BF16 | 320x10240 | 6.6 |
> | `mtp.layers.0.mlp_hyper_connection.input_mix_weight_up.weight` | BF16 | 10240x320 | 6.6 |
> | `mtp.layers.0.self_attn.indexer.index_qk_proj.weight` | BF16 | 640x2560 | 3.3 |
> | `mtp.layers.0.self_attn.indexer.k_layernorm.weight` | BF16 | 128 | 0.0 |
> | `mtp.layers.0.self_attn.indexer.q_layernorm.weight` | BF16 | 128 | 0.0 |
> | `mtp.layers.0.self_attn.k_norm.weight` | BF16 | 256 | 0.0 |
> | `mtp.layers.0.self_attn.k_proj.weight` | BF16 | 512x2560 | 2.6 |
> | `mtp.layers.0.self_attn.o_proj.weight` | BF16 | 2560x6144 | 31.5 |
> | `mtp.layers.0.self_attn.q_norm.weight` | BF16 | 256 | 0.0 |
> | `mtp.layers.0.self_attn.q_proj.weight` | BF16 | 12288x2560 | 62.9 |
> | `mtp.layers.0.self_attn.v_proj.weight` | BF16 | 512x2560 | 2.6 |
> | `mtp.pre_fc_norm_embedding.weight` | BF16 | 2560 | 0.0 |
> | `mtp.pre_fc_norm_hidden.weight` | BF16 | 10240 | 0.0 |
> 
> mtp.fc_embedding.weight 100%
> mtp.fc_hidden.weight 100%
> mtp.hyper_connection_mixer.hc_norm.weight 100%
> mtp.hyper_connection_mixer.input_mix_weight_down.weight 100%
> mtp.hyper_connection_mixer.input_mix_weight_up.weight 100%
> mtp.layers.0.attn_hyper_connection.block_inject_weight.weight 100%
> mtp.layers.0.attn_hyper_connection.hc_norm.weight 100%
> mtp.layers.0.attn_hyper_connection.input_mix_weight_down.weight 100%
> mtp.layers.0.attn_hyper_connection.input_mix_weight_up.weight 100%
> mtp.layers.0.mlp.experts.down_proj 4%
> mtp.layers.0.mlp.experts.down_proj 8%
> mtp.layers.0.mlp.experts.down_proj 12%
> mtp.layers.0.mlp.experts.down_proj 16%
> mtp.layers.0.mlp.experts.down_proj 20%
> mtp.layers.0.mlp.experts.down_proj 24%
> mtp.layers.0.mlp.experts.down_proj 28%
> mtp.layers.0.mlp.experts.down_proj 32%
> mtp.layers.0.mlp.experts.down_proj 36%
> mtp.layers.0.mlp.experts.down_proj 40%
> mtp.layers.0.mlp.experts.down_proj 44%
> mtp.layers.0.mlp.experts.down_proj 48%
> mtp.layers.0.mlp.experts.down_proj 52%
> mtp.layers.0.mlp.experts.down_proj 56%
> mtp.layers.0.mlp.experts.down_proj 60%
> mtp.layers.0.mlp.experts.down_proj 64%
> mtp.layers.0.mlp.experts.down_proj 68%
> mtp.layers.0.mlp.experts.down_proj 72%
> mtp.layers.0.mlp.experts.down_proj 76%
> mtp.layers.0.mlp.experts.down_proj 80%
> mtp.layers.0.mlp.experts.down_proj 84%
> mtp.layers.0.mlp.experts.down_proj 88%
> mtp.layers.0.mlp.experts.down_proj 92%
> mtp.layers.0.mlp.experts.down_proj 96%
> mtp.layers.0.mlp.experts.down_proj 100%
> mtp.layers.0.mlp.experts.gate_up_proj 2%
> mtp.layers.0.mlp.experts.gate_up_proj 4%
> mtp.layers.0.mlp.experts.gate_up_proj 6%
> mtp.layers.0.mlp.experts.gate_up_proj 8%
> mtp.layers.0.mlp.experts.gate_up_proj 10%
> mtp.layers.0.mlp.experts.gate_up_proj 12%
> mtp.layers.0.mlp.experts.gate_up_proj 14%
> mtp.layers.0.mlp.experts.gate_up_proj 16%
> mtp.layers.0.mlp.experts.gate_up_proj 18%
> mtp.layers.0.mlp.experts.gate_up_proj 20%
> mtp.layers.0.mlp.experts.gate_up_proj 22%
> mtp.layers.0.mlp.experts.gate_up_proj 24%
> mtp.layers.0.mlp.experts.gate_up_proj 26%
> mtp.layers.0.mlp.experts.gate_up_proj 28%
> mtp.layers.0.mlp.experts.gate_up_proj 30%
> mtp.layers.0.mlp.experts.gate_up_proj 32%
> mtp.layers.0.mlp.experts.gate_up_proj 34%
> mtp.layers.0.mlp.experts.gate_up_proj 36%
> mtp.layers.0.mlp.experts.gate_up_proj 38%
> mtp.layers.0.mlp.experts.gate_up_proj 40%
> mtp.layers.0.mlp.experts.gate_up_proj 42%
> mtp.layers.0.mlp.experts.gate_up_proj 44%
> mtp.layers.0.mlp.experts.gate_up_proj 46%
> mtp.layers.0.mlp.experts.gate_up_proj 48%
> mtp.layers.0.mlp.experts.gate_up_proj 50%
> mtp.layers.0.mlp.experts.gate_up_proj 52%
> mtp.layers.0.mlp.experts.gate_up_proj 54%
> mtp.layers.0.mlp.experts.gate_

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.