Issues / #1141
#1141 Windows AMD: prebuilt HIP zip running a model on a discrete card (Radeon AI PRO R9700, gfx1201)
open · @KongChengZhi · 0 コメント · GitHub で見る
BenchmarksSetup & installServer & APIAMD / HIPModels & quantsDocumentationWindowsLinux
本文
**TL;DR** — the ready-made Windows HIP zip runs a model on a discrete card here (Radeon AI PRO R9700, gfx1201, 32 GB): `strata-device --selftest` OK, IQ3_XXS at 131072 context in low-RAM resident mode, **50-81 tok/s** decode, **870 tok/s** prompt read, **zero** blob reads from the file while answering, >96% expert-cache hit rate.
Two findings, both A/B measured on this card:
1. **The PCIe probe picks `pcie_frac` too high on Windows.** It chose 0.55 from a 54.3 GB/s host->device reading; `--calibrate` measured 0.00 as **13% faster** (60.9 vs 53.8 tok/s), ~20% total together with `spec-min-p` 0.70. Consistent with #377 (the probe times copies on the host clock on Windows).
2. **WMMA is a net loss here.** +16% prompt read but **-18% decode**, against the +36-50% `docs/AMD_HIP.md` reports for this card (those numbers are from Linux).
---
This is the discrete-card Windows AMD run that `docs/AMD_HIP.md` asks for ("The ready-made zip itself has not run a model on a discrete card yet - please report"). It works; below are the numbers, plus one finding that looks like it matters beyond this card.
## Hardware / software
| | |
|---|---|
| Card | AMD Radeon AI PRO R9700, gfx1201, 31.9 GiB, wave32, 32 CUs |
| Driver | AMD Software 26.Q3, `32.0.31036.15` |
| OS | Windows 11 Pro |
| CPU | Ryzen 9 9950X (16C/32T, AVX-512) |
| RAM | 47.1 GB (2x24 GB DDR5-6000) |
| Engine | 0.1.39, repo commit `6f32ec070f23ced9f50e704d854d775da52591ab` |
| Model | Qwen3.8-Flash-Next **IQ3_XXS**, `--max-context 131072`, `--kv int8`, `--resident-experts` |
`setup --check` saw this as: *"fits in the low-RAM mode (the GPU holds ~63% of its experts, the rest stays in RAM)"* for IQ3_XXS and ~53% for IQ3_S.
## `strata-device --list-devices`
```
device 0: AMD Radeon AI PRO R9700
arch gfx1201, 31.9 GiB, wave32
```
## `strata-device --selftest`
```
strata: Windows budgets 29936 of this card's 32624 MiB for this process
device 0: AMD Radeon AI PRO R9700
HIP arch gfx1201 wave32 (compiled for gfx1100,gfx1101,gfx1102,gfx1200,gfx1201,gfx1030)
multiprocessors 32
VRAM total / free 31.859 GiB (34208743424 B) / 29.087 GiB (31232118754 B)
driver / runtime 71726391 / 71726391
HIP runtime D:\Workstation\Strata\engine\amdhip64_7.dll
memory plan --max-context 20480
KV + indexer keys 0.267387 GB
recurrent state 0.000000 GB (GDN recurrence + conv history, NOT evictable)
expert cache 5.675613 GB (4105 slots of 1,382,400 B)
VRAM pool 5.943000 GB
card free 29.087 GiB (31232118754 B)
plan + KV vs free FITS
selftest: arena 0.062 GiB (67108864 B), used 0.002 GiB (2097152 B) after two 1 MiB allocations
selftest: over-allocation refused as required
strata-device selftest OK
```
## End of `strata-iq3_xxs.log` (startup)
```
strata generate: PCIe probe: 54.3 GB/s host->device (best of 53.7 54.2 54.3 53.8) -> pcie_frac 0.55 (default 0.55)
strata generate: native pack: ...\packs\iq3_xxs experts (largest blob 2.33 MB), token embedding IQ3_S in mapped host memory (260 MiB)
strata generate: profile ...\data\expert-profile.bin: 24576 ranked pairs, built for 24576 slots
strata mtp: draft layer loaded, 950 MiB of VRAM (experts 675, dense 111), files read in 0.25 s (3182 MiB/s)
strata generate: GPU 0: AMD Radeon AI PRO R9700 (gfx1201)
strata: Windows budgets 31704 of this card's 32624 MiB for this process; free VRAM is counted within that
strata generate: expert cache auto: 24.28 GiB free, 700 MiB reserved (+184 MiB for the draft head) -> 10794 slots
strata generate: only 599 MiB free once the slots are written (reserve 700 MiB); shrinking the expert cache
strata generate: expert cache 13687 slots, 22.26 GiB of VRAM; policy is
PROFILE, ranked by routing frequency, no eviction.
strata generate: R4 hit path ON - resident experts are computed on the GPU
strata generate: 15 expert-pool workers + the host thread
strata generate: session is up (engine 0.1.39)
strata serve: prompt chunk auto: 8192 tokens, a 96-slot ring
prefill gemm: hipBLASLt tuning enabled (32 rows, gfx1201, version 100500)
FileExpertSource: resident 21.08 GiB, pinned 0.00 GiB
strata generate: resident RAM mode: 21.08 GiB of experts in RAM (pageable), 13687 in the GPU cache;
adaptive swaps exchange them with the GPU cache (no file reads)
strata serve: 393 MiB of VRAM free with everything loaded
```
The shipped `gfx1201-hipblaslt-100500.txt` table was picked up automatically and reports as enabled.
## Speed (from the server's own per-request lines)
```
strata serve: prompt 22 tokens = 0 reused + 22 read in 215 ms (102.5 tok/s), 128 generated in 2296 ms (55.8 tok/s), drafts accepted 73 of 95
strata serve: prompt 34 tokens = 0 reused + 34 read in 319 ms (106.6 tok/s), 128 generated in 1834 ms (69.8 tok/s), drafts accepted 91 of 98
strata serve: prompt 4876 tokens = 0 reused + 4876 read in 5601 ms (870.5 tok/s), 84 generated in 1039 ms (80.9 tok/s)
strata serve: decode expert cache hit rate: 96.4% (70769 hits / 73440 lookups)
strata serve: resident RAM: 20.81 GiB of experts in RAM, 5052 exchanged with the VRAM tier, 0 blob reads from the file
strata serve: expert tiers: GPU 70769 hits this request; since the start RAM 55875 blobs, files 0 blobs 24187.2 MB read
```
Decode runs **50-81 tok/s** depending on cache warmth (the first answer after a start is the slow one). The low-RAM resident mode does what it says: **zero blob reads from the file while answering** throughout a long session.
## Finding: on Windows the PCIe probe sets `pcie_frac` too high for this card
The startup probe measured **54.3 GB/s** host->device and chose `pcie_frac 0.55` (the default). `--calibrate` then measured decode speed across the range:
```
PCIe share 0.00: 60.9 tok/s
PCIe share 0.20: 58.0 tok/s
PCIe share 0.35: 54.0 tok/s
PCIe share 0.55: 53.8 tok/s <- what the probe chose
PCIe share 0.75: 55.9 tok/s
draft floor 0.30: 51.5 tok/s
draft floor 0.50: 57.1 tok/s <- previous default
draft floor 0.70: 61.6 tok/s
15 workers: 64.5 tok/s <- best; 10 workers 56.4, 8 workers 58.0
[ok] tuned for this PC: --pcie-frac 0.00, --spec-min-p 0.70 (64.5 tok/s)
```
So **0.00 is 13% faster than the 0.55 the probe picked**, and the two tuning changes together are worth about 20% over the previous defaults (53.8 -> 64.5 tok/s).
This looks consistent with #377's note that on Windows the PCIe probe times its copies on the host clock (HIP events read impossible speeds there), so the probe can overestimate the link and leave `pcie_frac` too high. If that is what is happening, **the probe may be unreliable on Windows in general** — the RDNA4 R9700 has a fast link and a strong CPU, so copying missed experts over PCIe loses to just letting the CPU compute them.
Practical consequence for anyone on Windows + AMD: **run `START-HERE.bat --calibrate` explicitly.** It is the only thing here that corrected an automatically chosen value. (Setup does not offer calibration on AMD by itself, per #566; 0.1.39 accepts a manual run and remembers the result for the card.)
## Finding: WMMA is a net loss on this card (prompt faster, decode slower)
`STRATA_HIP_WMMA=1` with `STRATA_PREFILL_RING=96` on gfx1201 + int8 KV, A/B with the same long prompt (unique first tokens each run so the KV prefix cache cannot be reused):
| | prompt read (4876 tok) | decode |
|---|---|---|
| `STRATA_HIP_WMMA=1`, ring 96 | 1011.5 tok/s | 61.1 tok/s |
| off | 871.0 tok/s | **74.2 tok/s** |
**+16% prompt, -18% decode.** The prompt side is well short of the +36-50% `docs/AMD_HIP.md` reports for this card (those numbers are from Linux runs). Since prompt read is paid once per request and decode once per token, this is a loss for chat and agent use, so I left WMMA off. The smaller `STRATA_PREFILL_RING=96` that goes with WMMA looks like the decode side of it.
## Notes and limits
- **Community measurements, not maintainer-validated.** Everything above comes from my own `strata-iq3_xxs.log` and `--calibrate` output.
- Single request at a time (server design), so numbers move with expert-cache warmth; I used the engine's own lines rather than wall-clock timing.
- **Images are off** on Windows AMD as documented; this run is text-only.
- Context was kept at 131072. gfx1201 long-context behaviour beyond 16K is still unexercised.
- `--calibrate`'s numbers come from its own harness and are not directly comparable to the per-request lines above, but they agree in band.
- `strata-device --selftest` was run while the server was up (29 GiB still free) and passed, including the over-allocation refusal.
Happy to re-run anything, or to try IQ3_S on this card if that would help.関連リンク
インストール・モデル・リリースへの站内リンク。