Pull requests / #1078
#1078 bench: RX 6800M (gfx1031) community report - self-built HIP engine on Windows
closed · @Hugua700 · 0 comments · View on GitHub
BenchmarksAMD / HIPModels & quantsDocumentationWindows
Description
Results-only community benchmark report, as `docs/COMMUNITY_BENCHMARKS.md` asks for. No engine changes. **Hardware:** AMD Radeon RX 6800M (gfx1031, Navi 22, 12 GB) — one mobile card — Ryzen 9 5900HX, 31.4 GB RAM, Windows 11 (build 26300), AMD driver 32.0.21045.5002. **This is the first gfx1031 report and the smallest configuration in `bench/results/`.** **Software:** Strata `1735d64` (v0.1.40), HIP engine **built from source** with `STRATA_HIP_ARCHS=gfx1031` (the ready-made Windows HIP zip has no gfx1031 code), ROCm 10.2.0a20260930 from the TheRock wheels. CPU image encoder built too. **Model:** `SC117/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-GGUF`, **Q2_0**, shard 1 sha256 checked against the Hugging Face LFS id; pack from `tools/iq_pack.py`, `data/expert-profile.bin`, MTP draft layer. Context 65,536, KV int8 with `--kv-resident 32768`, `--resident-experts`. **What was measured:** three prompt sizes (556 / 2,309 / 8,726 tokens) × 3 runs each, at the default settings and with `STRATA_HIP_PROMPT_F16=1`; plus six recall checks. One request at a time, greedy, 256-token output cap, model load excluded, warm-up discarded. Prefix reuse is defeated on purpose with a per-run nonce line so every run reads the full prompt (`timings.cache_n = 0` everywhere). **Headline numbers:** prefill **50–114 tok/s**, decode **9–27 tok/s**, TTFT 8.6–94.6 s. Far below the published figures, for two reasons the engine states itself: the experts do not fit in RAM (`resident complement 26.46 GiB exceeds available RAM (16.66 GiB) … the rest are read from the model folder`), and this CPU has no AVX-512 so the expert kernels run AVX-2. **The one result worth a maintainer's attention:** the FP16 prompt switch gives **+6 % to +29 %** here, against **+69 % to +101 %** in `bench/results/2026-10-04-rdna2-fp16-prompt` (RX 6900 XT, 128 GB RAM, model fits). Both are the same switch, and the difference looks consistent rather than contradictory: when the experts stream from disk, prefill is not bound by the prompt GEMM, so the GEMM win sits on top of a much lower floor. The report also treats the small decode deltas as run-to-run noise, matching that report's finding that decode does not use these GEMMs. The report also records two notes for other low-RAM AMD users: the abliterated repository has no IQ1_M quant (its ladder starts at Q2_0, which `--check` correctly reports as not fitting 32 GB), and `--expert-profile` follows the model family (Coder is 48×256, the plain model 48×512), so it must be switched together with the weights. **Limitations** (also in the report): longest prompt 8,727 tokens against a configured 65,536 — no 32K/128K sweep; one model and one quantization; decode includes the MTP/suffix draft path; RAM saturated at 0.1–0.2 GB free during the runs; "Performance" power plan. The machine has a documented power-delivery fault (four spontaneous power-offs under sustained load), so the measurements were deliberately taken in short segments rather than one long sweep — that is also why there is no long-context point.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.