Pull requests / #1565
#1565 bench: community report — 2x Radeon RX 9070 (gfx1201), IQ3_S @ 262144 on ROCm
open · @eldiaboloz · 0 Kommentare · Auf GitHub
BenchmarksMulti-GPUAMD / HIPModels & quantsDocumentationLinux
Beschreibung
# bench: community report — 2x Radeon RX 9070 (gfx1201), IQ3_S @ 262144 on ROCm Results-only community benchmark. No engine changes. ## Summary One of the first RDNA4 / Linux reports: **IQ3_S at the model's native 262,144-token context** with the layer split across **2x AMD Radeon RX 9070** (RX 9070 + RX 9070 XT, gfx1201) on ROCm 7.2.4, engine 0.1.40.2. All speed rows are freshly processed (no prefix reuse) unless stated. Median: **841.6 prefill / 67.1 decode tok/s at 4,226 prompt tokens**, **1,488.6 / 64.5 at 33,428**, **1,916.4 / 60.0 at 130,542**, and one run at **248,708 tokens: 1,844.8 / 59.5**. Recall: 6/6 needles. Report folder: `bench/results/2026-10-08-community-2x-rx9070-iq3s/` ## Hardware - 2x AMD Radeon RX 9070, 16 GB each, gfx1201 (Navi 48): **GPU 0 = RX 9070** (220 W stock), **GPU 1 = RX 9070 XT capped to 225 W** (default 304 W). Both at **PCIe 4.0 x8**; engine PCIe probe 14.4 GB/s per card (`amdgpu_top` DPM range Gen1x8–Gen4x8). - AMD Ryzen 9 5950X (16C/32T, AVX2, no AVX-512); 15 expert-pool workers. 126 GiB DDR4. NVMe. 1000 W PSU. ## Software - Arch Linux, kernel 7.2.7-arch1-1, amdgpu. - ROCm 7.2.4, hipBLASLt 1.2.2 (`hipblaslt 7.2.4-1`), HIPBLASLt enabled with `tools/hip/gfx1201-hipblaslt-100202.txt`. - Strata commit `e8ca9af` (engine 0.1.40.2), local HIP gfx1201 source build. ## Model and configuration - `ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF`, **IQ3_S** (two shards; shard 2 is the PLE table). - Native pack (`tools/iq_pack.py`), bundled expert profile, MTP draft layer, no vision. - Context **262,144** (native), `--kv int8 --kv-resident 32768`, `"gpu": [1,0]`, `layer_split "24"`. - Expert cache `auto`: **10,790 slots / 20,869 MiB** total. `--prefill auto:32768` chooses a **19,456-token chunk** (capped by the second card's slots). `--pcie-frac` probed to **0.40**. - `--ple-io ram` (PLE locked in RAM, 26.8 GiB), `--remote-expert-opt`, conversation cache 8 GiB/4, `--spec 4 --spec-min-p 0.5`. Requests set temperature 0 and reasoning off (server sampling defaults: temp 0.6, top_p 0.95, top_k 20, rep 1.05, penalty_last_n 256). ## Results (median [min-max], 3 runs; 260-token cap) | Prompt tokens | Reused | Prefill tok/s | Decode tok/s | | ---: | ---: | --- | --- | | 4,226 | 0 | 841.6 [814.8–841.9] | 67.1 [65.7–70.6] | | 33,428 | 0 | 1,488.6 [1,469.4–1,492.5] | 64.5 [63.4–67.9] | | 130,542 | 0 | 1,916.4 [1,913.1–1,920.9] | 60.0 [56.5–62.0] | | 248,708 (single) | 0 | 1,844.8 | 59.5 | - Warm reuse: 8,217/8,224 tokens reused; the same 8K prompt went 12.0 s → 3.86 s. - TTFT (streaming, short prompt): **1.81 s [1.80–1.81]**. - Memory: engine RSS **86.6–89.8 GiB** (26.8 GiB locked PLE), `MemAvailable` ≥ 29.3 GiB, VRAM 15.59/15.69 GiB. No paging, OOM, failed or cancelled requests. ## Correctness `tools/needle_bench.py` found **6/6** needles at depths 10/50/90% at both 32K and 128K. ## Limitations - One machine, one quantization, one configuration, one synthetic prompt family, greedy decoding. - Server sampling defaults were overridden per request; `--ple-io ram` is a deliberate host choice. - The 248,708-token case is a single run. Agentic/tool use, coding correctness, vision, concurrency and sustained thermal runs were not measured (cards are power-capped). ## Files `README.md`, `benchmark.py`, `results.json`, `summary.json`, `ttft.json`, `needles.json`, `telemetry.jsonl`, `memory-summary.json`, `model-provenance.json`, `mtp-manifest.json`, `initial-status.json`, `final-status.json`, `engine.log`, `strata-iq3_s.json` (API key redacted), `BUILD.json`. No credentials, models or packs are included. ## Extra notes - Same machine as a separate PCIe-link A/B (both cards Gen3 x8 vs Gen4 x8); that and an OCuLink x8+x4 run are planned as a follow-up.
Mehr auf der Site
Links zu Install, Modellen, Releases.