贡献 / #1655
#1655 bench: community report, 4x RX 7900 XT (gfx1100), Flash-Next IQ3_S at 262K, five opt-ins off vs on, engine 0.1.41
open · @Cass67 · 0 评论 · 去 GitHub 看
BenchmarksAMD / HIPModels & quantsDocumentationWindowsLinux
说明
Results-only submission, no engine changes. Adds bench/results/2026-10-09-community-4x-rx-7900-xt/ following docs/COMMUNITY_BENCHMARKS.md. Hardware: 4x RX 7900 XT 20 GiB (gfx1100), Core i9-9900K, 60.4 GiB RAM, ASUS WS Z390 PRO. Three cards behind a PLX PEX 8747 at Gen3 x16/x8/x8 (one shared Gen3 x16 uplink), the fourth on the chipset at Gen3 x4. Linux 7.0.0-34, ROCm 10.0.0. Software: Strata v0.1.41 (fb58e0d), clean tree, source build in a container (HIP, `STRATA_PREFILL_MMQ`, `STRATA_MMQ_KQUANTS`, gfx1100). Model: ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF IQ3_S (both shards match revision ed59f92 by size and LFS sha256), 262,144 context, int8 KV, layer split 12,24,36, every expert in VRAM, MTP `--spec 4 --spec-min-p 0.5`. Configurations tested: two arms differing only in five documented opt-ins (`STRATA_SPLIT_OWN`, `STRATA_STAGE_TRIM`, `STRATA_PF_FUSED`, `STRATA_PF_GEMM`, `STRATA_PF_SWITCH_MIN_T=4096`), each a fresh process: one cold first request, two warm-ups, then 3 runs each at 4,096 / 32,768 / 128,000 prompt tokens (sized with the pack's tokenizer from repository text, unique first line, cache_n 0 on all), 256-token cap, greedy, reasoning off; 6 needle checks per arm (12 of 12 found). VRAM per card sampled every second. Headline (medians, opt-ins on / off): prompt 2,052 / 756, 3,804 / 1,722, 4,307 / 2,433 tok/s; decode 60.0 / 59.3, 52.7 / 52.0, 50.4 / 49.6 tok/s at 4K / 32K / 128K. Peak VRAM 62.6 / 69.0 GiB over the four cards. Limitations: one model, one machine, 3 runs per cell, answers cut at 256 tokens and not graded. The chipset-attached card reports a 0 W power cap. Pack and MTP preparation commands were not recorded (hashes are in the README). Context: a four-stage `--pipeline-windows` change measured on this machine is #1656 (not used for these results).
本站相关内容
相关页面的快捷入口。