Issues / #953

#953 Windows AMD report: UD-IQ4_XS on RX 7900 XTX (gfx1100) works, 34-64 tok/s decode

open · @Huobi-cloud · 2 commentaires · Sur GitHub

BenchmarksSetup & installServer & APIAMD / HIPModels & quantsWindows

Description

Windows AMD report: UD-IQ4_XS on RX 7900 XTX (gfx1100) works, 34-64 tok/s decode

### Setup
- GPU: AMD Radeon RX 7900 XTX 24 GB (gfx1100), plus the Ryzen iGPU (gfx1036, correctly ignored)
- CPU: AMD Ryzen 7 9800X3D (AVX-512), RAM: 64 GB DDR5 (setup reports 62 GB)
- OS: Windows 11, build 10.0.26300.9457
- Driver: AMD Software: Adrenalin Edition 26.8.1 (released 2026-08-17)
- Strata engine 0.1.39, fresh git clone on 2026-10-05, default settings
- Install: START-HERE.bat --yes --family unsloth --model UD-IQ4_XS --gpu 1 --no-start

### Notes from setup
- --check listed UD-IQ4_XS as "NVIDIA only so far, untested on AMD" - it installed and runs fine on AMD.
- Warning: no hipBLASLt tuning table for gfx1100 with hipBLASLt 100500 (only 100100/100200 shipped), so plain hipBLAS is used for the prompt's dense GEMMs.

### Results
Web UI, a new chat per prompt, default settings (STRATA_SH_STREAM not set). Token counts and speeds as shown by the web UI.

| Prompt | Tokens | Speed |
|---|---|---|
| "Explain in detail how a concrete foundation is poured for a family house, step by step, in about 500 words." | 9,304 | **64.4 tok/s** |
| Romanian: a 400-word guide on choosing a gas boiler for a 150 m² house | 921 | **49.0 tok/s** |
| "Write a Python script that reads a CSV file with columns date, product, quantity, price and prints the total sales per month, sorted by date." | 1,219 | **54.1 tok/s** |
| "A truck carries 12 tons of sand. It unloads 1/3 at the first site and 2.5 tons at the second. How much is left? Explain briefly." (answer correct: 5.5 t) | 173 | **44.7 tok/s** |

First two short requests right after the first start (from the log):

| Request | Prompt | Decode | MTP drafts accepted | Expert cache hit rate |
|---|---|---|---|---|
| 1 | 58 tokens read, 44.8 tok/s | 102 tokens, 34.7 tok/s | 60 / 86 | 69.5% |
| 2 | 82 new tokens read (53 reused), 39.4 tok/s | 359 tokens, 33.7 tok/s | 177 / 355 | 90.1% |

- Decode speed rises with longer answers and as the expert cache warms up (69.5% -> 90.1% hit rate between the first two requests).
- RAM tier: 38.00 GiB of experts page-locked in RAM; 6170 expert slots (13.90 GiB) in VRAM; 560 MiB VRAM free with everything loaded.
- File tier: 17.4 GiB of experts read from the GGUF in place.
- Answers were coherent and correct in English and Romanian. Long prompts (20K+) not tested yet.

### strata-device
```
device 0: AMD Radeon RX 7900 XTX
  arch gfx1100, 24.0 GiB, wave32
device 1: AMD Radeon(TM) Graphics
  arch gfx1036, 43.6 GiB, wave32
  cannot run: GPU 1 (gfx1036) is not an architecture this Strata engine was compiled for

strata: Windows budgets 23749 of this card's 24560 MiB for this process
device 0: AMD Radeon RX 7900 XTX
  HIP arch           gfx1100 wave32 (compiled for gfx1100,gfx1101,gfx1102,gfx1200,gfx1201,gfx1030)
  multiprocessors    48
  VRAM total / free  23.984 GiB / 23.047 GiB
  driver / runtime   71726391 / 71726391
  HIP runtime        D:\Strata\engine\amdhip64_7.dll
memory plan --max-context 20480 ... plan + KV vs free  FITS
strata-device selftest OK
```

### Log excerpt (strata-unsloth-ud-iq4_xs.log)
```
strata generate: PCIe probe: 28.2 GB/s host->device -> pcie_frac 0.55 (default 0.55)
strata generate: GPU 0: AMD Radeon RX 7900 XTX (gfx1100)
strata generate: experts via mmap (--mmap-experts; the GGUF shards in place, no experts.bin)
strata generate: expert cache auto: 16.81 GiB free, 700 MiB reserved (+218 MiB for the draft head) -> 4906 slots
strata generate: only 0 MiB free once the slots are written (reserve 700 MiB); shrinking the expert cache
strata generate: only 0 MiB free once the slots are written (reserve 700 MiB); shrinking the expert cache
strata generate: expert cache 6170 slots, 13.90 GiB of VRAM
strata generate: session is up (engine 0.1.39)
FileExpertSource: RAM budget 38.00 GiB: 16836 of the 41.53 GiB of experts the GPU cache does not hold, by profile rank; the rest are read from the files
strata generate: resident RAM mode: 38.00 GiB of experts in RAM (page-locked), 6170 in the GPU cache
strata serve: 560 MiB of VRAM free with everything loaded
strata serve: prompt 58 tokens = 0 reused + 58 read in 1295 ms (44.8 tok/s), 102 generated in 2936 ms (34.7 tok/s), drafts accepted 60 of 86
strata serve: decode expert cache hit rate: 69.5% (36060 hits / 51915 lookups)
strata serve: prompt 135 tokens = 53 reused + 82 read in 2081 ms (39.4 tok/s), 359 generated in 10665 ms (33.7 tok/s), drafts accepted 177 of 355
strata serve: decode expert cache hit rate: 90.1% (219527 hits / 243652 lookups)
```

Full log available on request. Thanks for Strata!

Sur le site

Liens install, modèles, releases.