Issues / #647

#647 Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS works on double Vega 20 GPU (Radeon Pro VII 16gb each).

closed · @Marczellito · 1 Kommentare · Auf GitHub

Setup & installMulti-GPUAMD / HIPModels & quantsLinux

Beschreibung

My hardware and software setup:
GPU: 32gb, 2x AMD Radeon Pro VII (16gb VRAM each) (= gfx906, = Vega 20)
RAM: 64 gb RAM (8x 8gb DDR4-2133 ECC) [probably running at 1866 MHz though]
CPU: Intel Xeon E5-2696 v3
OS storage: Samsung 980 NVMe
OS: CachyOS Linux (recent stable official version, not modified)
Model: Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS

Logs:
[strata-swift-iq3_xxs.log](https://github.com/user-attachments/files/33004465/strata-swift-iq3_xxs.log)

I've tried running it in Unsloth before, so Strata's MTP is in different folder than the model itself (which is stored on external hdd).
Two problems as a layman: 
1. when idle model keeps unloading from memory, default option should be to keep it in memory until other specified, because reloading model takes relatively long time (~5-7minutes).
2. the time to abort prefill is set too short - from what i understand in the logs - hardcoded time-out prevented the model from analyzing even short .txt file before answering (at least in my setup).

I used Unsloth Studio to set it all up with unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q6_K_XL 
This is a summary that it produced during setup procedure (pasted from AI, minimal editing, SLOP WARNING):

WHAT WORKED:
✓ ROCm 7.2.4 installed and working on CachyOS
✓ Both GPUs detected as gfx906 by rocminfo
✓ Strata engine COMPILED successfully for gfx906 architecture
✓ Swift 1.5 IQ3_XXS model files recognized and validated
✓ Model prepared (tokenizer, experts.bin, pack created)
✓ MTP draft layer downloaded and verified
✓ Layer split across both GPUs detected correctly
✓ Expert arena loaded (~40 GB into RAM at 0.11 GB/s)
✓ KV cache streaming configured (32768/65536 cells in VRAM)
✓ PCIe bandwidth probed: 13.8 GB/s between GPUs
✓ MTP draft layer loaded into VRAM (835 MiB)

WHAT DID NOT WORK / KNOWN ISSUES:
1. First full startup attempt crashed during expert cache initialization
   - Engine loaded all weights successfully
   - Crashed at: "expert cache auto: 12.13 GiB free, 700 MiB reserved -> 5275 slots"
   - Log shows no explicit error message (silent crash)
   - Likely cause: VRAM exhaustion trying to register large memory regions on gfx906
2. No hipBLASLt tuning table for gfx906
   - Using plain hipBLAS for prompt dense matrix products (slower prompts)
   - Does NOT affect model quality, only prompt processing speed
3. Vega 20 (gfx906) is not officially supported
   - Strata kernels are designed for wave32; Vega uses wave64
   - Patched to allow wave64 but behavior may be suboptimal
   - No validated performance data exists for this architecture

Thank you for your project, i hope my logs will be somehow usefull!

Mehr auf der Site

Links zu Install, Modellen, Releases.