Issues / #942

#942 AMD `gfx1102 / RX 7600 XT` - successful real model run on Strata 0.1.39

open · @Lefox-DeMod · 2 评论 · 在 GitHub 查看

BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationLinux

描述

# gfx1102 / RX 7600 XT — successful model run on Strata 0.1.39

Measured on 2026-10-05 by me

`docs/AMD_HIP.md` currently notes that **gfx1102 builds but has not yet had a reported model run**. I tested it on a Radeon RX 7600 XT and successfully ran a large Qwen3.8-Flash-Next model for four sessions.

This is submitted as an issue rather than a benchmark PR because the measurements below come from normal interactive usage rather than the standardized benchmark procedure.

## Hardware

* **GPU:** AMD Radeon RX 7600 XT 16 GB, Navi 33, `gfx1102`, PCI ID `0x7480`
* **CPU:** Intel Core i7-13700K, 16C/24T, AVX2
* **RAM:** 64 GB
* **PCIe:** PCIe 4.0, negotiated `16.0 GT/s x8`
* **Host OS:** Void Linux, kernel `6.18.54_1`
* **Driver:** kernel `amdgpu`

## Build environment

Strata was built and tested inside a **Distrobox container based on a current Arch Linux userspace**, running on the Void Linux host.

The container provided the ROCm/HIP toolchain:

* ROCm 7.2.x
* ROCm Clang 22.0.0
* CMake 4.4.3
* Ninja 1.13.2
* GNU C/C++ 16.2.1

The host Void Linux installation does not have a system ROCm installation.

Strata:

* **Version:** `0.1.39`
* **Commit:** `d63ddbc6a262e39e`
* **Backend:** HIP
* **llama.cpp:** `3cf0325`

## Build change

The only source change was adding `gfx1102` to the architecture gate in `setup.py`:

```diff
- AMD_ARCHS = ("gfx1100", "gfx1101", "gfx1200", "gfx1201", "gfx1030", "gfx1031")
+ AMD_ARCHS = ("gfx1100", "gfx1101", "gfx1102", "gfx1200", "gfx1201", "gfx1030", "gfx1031")
```

The normal `./setup.sh` build was then run with:

```text
-DSTRATA_ENABLE_HIP=ON
-DSTRATA_ENABLE_CUDA=OFF
-DSTRATA_BUILD_TESTS=OFF
-DSTRATA_PREFILL_MMQ=ON
-DCMAKE_HIP_ARCHITECTURES=gfx1102
```

CMake produced the expected validation warning:

```text
Strata HIP: gfx1102 builds, but it is not validated on a real card yet; please report results
```

Build result:

```text
130/130 targets
0 errors
```

No gfx1102-specific compiler or ISA errors were observed.

## Model

**Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS**

Split GGUF:

* `...00001-of-00002.gguf` — ~39.8 GB
* `...00002-of-00002.gguf` — ~36.2 GB

Strata configuration:

```text
Context:        65536
KV cache:       int8
KV resident:    32768
Expert cache:   auto
Prefill:        auto
MTP:            --spec 4 --spec-min-p 0.5
Reasoning:      none
Temperature:    0.6
top_p:          0.95
top_k:          20
```

The model used a locally prepared Strata native pack, expert profile, and MTP runtime.

## Test method

The model was served through Strata's OpenAI-compatible API on localhost.

Across four engine sessions:

* **146 requests**
* **38,490 generated tokens**
* prompts up to **51,543 tokens**
* prefix reuse up to approximately **51K tokens**

The runs were normal interactive usage rather than controlled benchmark repetitions, so the throughput values below are observed results.

## Results

### Decode throughput

| Workload                        | Runs |     Median |     Range |
| ------------------------------- | ---: | ---------: | --------: |
| Long outputs (1024–4096 tokens) |    6 | 37.8 tok/s | 33.6–38.7 |
| Medium outputs (256–512 tokens) |   20 | 42.7 tok/s | 28.0–49.6 |
| Warm follow-up requests         |  120 | 43.3 tok/s | 27.4–49.6 |

### Fresh prompt processing

|    Prompt size | Runs |      Median |       Range |
| -------------: | ---: | ----------: | ----------: |
|   1K–5K tokens |   18 | 252.3 tok/s | 194.2–280.3 |
|  5K–20K tokens |   11 | 297.4 tok/s | 213.4–310.2 |
| 20K–40K tokens |    5 | 305.0 tok/s | 297.7–309.1 |

Examples:

```text
21,680 prompt tokens -> 306.8 tok/s
3,825 generated tokens -> 38.2 tok/s

24,443 prompt tokens -> 309.1 tok/s
17,514 prompt tokens -> 301.1 tok/s
```

## Stability and observations

The four sessions completed without failed requests, GPU-side errors, or OOM conditions.

Other observed values:

* ~624 MiB free VRAM at steady state
* median expert-cache hit rate: **82.7%**
* overall MTP draft acceptance: **76.8%**
* PCIe host→device transfer measured by the engine: **13.4 GB/s** at best

There is currently no gfx1102-specific hipBLASLt tuning table, so the dense prefill path used plain hipBLAS.

## Limitations

Not tested:

* vision
* contexts beyond ~51K tokens
* concurrency
* sustained thermal load
* gfx1102 hipBLASLt tuning
* standardized `COMMUNITY_BENCHMARKS.md` runs
* formal quality/needle/coding benchmarks

This report covers one GPU, one model quantization, and one configuration.

## Suggested follow-up

* Add `gfx1102` to `AMD_ARCHS` so the local build patch is no longer required.
* Update `docs/AMD_HIP.md` to record the successful model run.
* Consider adding a gfx1102 hipBLASLt tuning table later.

Full build and engine logs can be provided for further investigation.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。