Issues / #942
#942 AMD `gfx1102 / RX 7600 XT` - successful real model run on Strata 0.1.39
open · @Lefox-DeMod · 2 comentários · No GitHub
BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationLinux
Descrição
# gfx1102 / RX 7600 XT — successful model run on Strata 0.1.39
Measured on 2026-10-05 by me
`docs/AMD_HIP.md` currently notes that **gfx1102 builds but has not yet had a reported model run**. I tested it on a Radeon RX 7600 XT and successfully ran a large Qwen3.8-Flash-Next model for four sessions.
This is submitted as an issue rather than a benchmark PR because the measurements below come from normal interactive usage rather than the standardized benchmark procedure.
## Hardware
* **GPU:** AMD Radeon RX 7600 XT 16 GB, Navi 33, `gfx1102`, PCI ID `0x7480`
* **CPU:** Intel Core i7-13700K, 16C/24T, AVX2
* **RAM:** 64 GB
* **PCIe:** PCIe 4.0, negotiated `16.0 GT/s x8`
* **Host OS:** Void Linux, kernel `6.18.54_1`
* **Driver:** kernel `amdgpu`
## Build environment
Strata was built and tested inside a **Distrobox container based on a current Arch Linux userspace**, running on the Void Linux host.
The container provided the ROCm/HIP toolchain:
* ROCm 7.2.x
* ROCm Clang 22.0.0
* CMake 4.4.3
* Ninja 1.13.2
* GNU C/C++ 16.2.1
The host Void Linux installation does not have a system ROCm installation.
Strata:
* **Version:** `0.1.39`
* **Commit:** `d63ddbc6a262e39e`
* **Backend:** HIP
* **llama.cpp:** `3cf0325`
## Build change
The only source change was adding `gfx1102` to the architecture gate in `setup.py`:
```diff
- AMD_ARCHS = ("gfx1100", "gfx1101", "gfx1200", "gfx1201", "gfx1030", "gfx1031")
+ AMD_ARCHS = ("gfx1100", "gfx1101", "gfx1102", "gfx1200", "gfx1201", "gfx1030", "gfx1031")
```
The normal `./setup.sh` build was then run with:
```text
-DSTRATA_ENABLE_HIP=ON
-DSTRATA_ENABLE_CUDA=OFF
-DSTRATA_BUILD_TESTS=OFF
-DSTRATA_PREFILL_MMQ=ON
-DCMAKE_HIP_ARCHITECTURES=gfx1102
```
CMake produced the expected validation warning:
```text
Strata HIP: gfx1102 builds, but it is not validated on a real card yet; please report results
```
Build result:
```text
130/130 targets
0 errors
```
No gfx1102-specific compiler or ISA errors were observed.
## Model
**Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS**
Split GGUF:
* `...00001-of-00002.gguf` — ~39.8 GB
* `...00002-of-00002.gguf` — ~36.2 GB
Strata configuration:
```text
Context: 65536
KV cache: int8
KV resident: 32768
Expert cache: auto
Prefill: auto
MTP: --spec 4 --spec-min-p 0.5
Reasoning: none
Temperature: 0.6
top_p: 0.95
top_k: 20
```
The model used a locally prepared Strata native pack, expert profile, and MTP runtime.
## Test method
The model was served through Strata's OpenAI-compatible API on localhost.
Across four engine sessions:
* **146 requests**
* **38,490 generated tokens**
* prompts up to **51,543 tokens**
* prefix reuse up to approximately **51K tokens**
The runs were normal interactive usage rather than controlled benchmark repetitions, so the throughput values below are observed results.
## Results
### Decode throughput
| Workload | Runs | Median | Range |
| ------------------------------- | ---: | ---------: | --------: |
| Long outputs (1024–4096 tokens) | 6 | 37.8 tok/s | 33.6–38.7 |
| Medium outputs (256–512 tokens) | 20 | 42.7 tok/s | 28.0–49.6 |
| Warm follow-up requests | 120 | 43.3 tok/s | 27.4–49.6 |
### Fresh prompt processing
| Prompt size | Runs | Median | Range |
| -------------: | ---: | ----------: | ----------: |
| 1K–5K tokens | 18 | 252.3 tok/s | 194.2–280.3 |
| 5K–20K tokens | 11 | 297.4 tok/s | 213.4–310.2 |
| 20K–40K tokens | 5 | 305.0 tok/s | 297.7–309.1 |
Examples:
```text
21,680 prompt tokens -> 306.8 tok/s
3,825 generated tokens -> 38.2 tok/s
24,443 prompt tokens -> 309.1 tok/s
17,514 prompt tokens -> 301.1 tok/s
```
## Stability and observations
The four sessions completed without failed requests, GPU-side errors, or OOM conditions.
Other observed values:
* ~624 MiB free VRAM at steady state
* median expert-cache hit rate: **82.7%**
* overall MTP draft acceptance: **76.8%**
* PCIe host→device transfer measured by the engine: **13.4 GB/s** at best
There is currently no gfx1102-specific hipBLASLt tuning table, so the dense prefill path used plain hipBLAS.
## Limitations
Not tested:
* vision
* contexts beyond ~51K tokens
* concurrency
* sustained thermal load
* gfx1102 hipBLASLt tuning
* standardized `COMMUNITY_BENCHMARKS.md` runs
* formal quality/needle/coding benchmarks
This report covers one GPU, one model quantization, and one configuration.
## Suggested follow-up
* Add `gfx1102` to `AMD_ARCHS` so the local build patch is no longer required.
* Update `docs/AMD_HIP.md` to record the successful model run.
* Consider adding a gfx1102 hipBLASLt tuning table later.
Full build and engine logs can be provided for further investigation.
No site
Links install, modelos, releases.