Pull requests / #1565

#1565 bench: community report — 2x Radeon RX 9070 (gfx1201), IQ3_S @ 262144 on ROCm

open · @eldiaboloz · 0 comments · View on GitHub

BenchmarksMulti-GPUAMD / HIPModels & quantsDocumentationLinux

Description

# bench: community report — 2x Radeon RX 9070 (gfx1201), IQ3_S @ 262144 on ROCm

Results-only community benchmark. No engine changes.

## Summary

One of the first RDNA4 / Linux reports: **IQ3_S at the model's native 262,144-token context** with
the layer split across **2x AMD Radeon RX 9070** (RX 9070 + RX 9070 XT, gfx1201) on ROCm 7.2.4,
engine 0.1.40.2. All speed rows are freshly processed (no prefix reuse) unless stated.

Median: **841.6 prefill / 67.1 decode tok/s at 4,226 prompt tokens**, **1,488.6 / 64.5 at 33,428**,
**1,916.4 / 60.0 at 130,542**, and one run at **248,708 tokens: 1,844.8 / 59.5**. Recall: 6/6 needles.

Report folder: `bench/results/2026-10-08-community-2x-rx9070-iq3s/`

## Hardware

- 2x AMD Radeon RX 9070, 16 GB each, gfx1201 (Navi 48): **GPU 0 = RX 9070** (220 W stock), **GPU 1
  = RX 9070 XT capped to 225 W** (default 304 W). Both at **PCIe 4.0 x8**; engine PCIe probe 14.4
  GB/s per card (`amdgpu_top` DPM range Gen1x8–Gen4x8).
- AMD Ryzen 9 5950X (16C/32T, AVX2, no AVX-512); 15 expert-pool workers. 126 GiB DDR4. NVMe. 1000 W PSU.

## Software

- Arch Linux, kernel 7.2.7-arch1-1, amdgpu.
- ROCm 7.2.4, hipBLASLt 1.2.2 (`hipblaslt 7.2.4-1`), HIPBLASLt enabled with
  `tools/hip/gfx1201-hipblaslt-100202.txt`.
- Strata commit `e8ca9af` (engine 0.1.40.2), local HIP gfx1201 source build.

## Model and configuration

- `ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF`, **IQ3_S** (two shards; shard 2 is the PLE table).
- Native pack (`tools/iq_pack.py`), bundled expert profile, MTP draft layer, no vision.
- Context **262,144** (native), `--kv int8 --kv-resident 32768`, `"gpu": [1,0]`, `layer_split "24"`.
- Expert cache `auto`: **10,790 slots / 20,869 MiB** total. `--prefill auto:32768` chooses a
  **19,456-token chunk** (capped by the second card's slots). `--pcie-frac` probed to **0.40**.
- `--ple-io ram` (PLE locked in RAM, 26.8 GiB), `--remote-expert-opt`, conversation cache 8 GiB/4,
  `--spec 4 --spec-min-p 0.5`. Requests set temperature 0 and reasoning off (server sampling
  defaults: temp 0.6, top_p 0.95, top_k 20, rep 1.05, penalty_last_n 256).

## Results (median [min-max], 3 runs; 260-token cap)

| Prompt tokens | Reused | Prefill tok/s | Decode tok/s |
| ---: | ---: | --- | --- |
| 4,226 | 0 | 841.6 [814.8–841.9] | 67.1 [65.7–70.6] |
| 33,428 | 0 | 1,488.6 [1,469.4–1,492.5] | 64.5 [63.4–67.9] |
| 130,542 | 0 | 1,916.4 [1,913.1–1,920.9] | 60.0 [56.5–62.0] |
| 248,708 (single) | 0 | 1,844.8 | 59.5 |

- Warm reuse: 8,217/8,224 tokens reused; the same 8K prompt went 12.0 s → 3.86 s.
- TTFT (streaming, short prompt): **1.81 s [1.80–1.81]**.
- Memory: engine RSS **86.6–89.8 GiB** (26.8 GiB locked PLE), `MemAvailable` ≥ 29.3 GiB, VRAM
  15.59/15.69 GiB. No paging, OOM, failed or cancelled requests.

## Correctness

`tools/needle_bench.py` found **6/6** needles at depths 10/50/90% at both 32K and 128K.

## Limitations

- One machine, one quantization, one configuration, one synthetic prompt family, greedy decoding.
- Server sampling defaults were overridden per request; `--ple-io ram` is a deliberate host choice.
- The 248,708-token case is a single run. Agentic/tool use, coding correctness, vision, concurrency
  and sustained thermal runs were not measured (cards are power-capped).

## Files

`README.md`, `benchmark.py`, `results.json`, `summary.json`, `ttft.json`, `needles.json`,
`telemetry.jsonl`, `memory-summary.json`, `model-provenance.json`, `mtp-manifest.json`,
`initial-status.json`, `final-status.json`, `engine.log`, `strata-iq3_s.json` (API key redacted),
`BUILD.json`. No credentials, models or packs are included.

## Extra notes

- Same machine as a separate PCIe-link A/B (both cards Gen3 x8 vs Gen4 x8); that and an
  OCuLink x8+x4 run are planned as a follow-up.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.