Issues / #590
#590 RX 7900 XTX (gfx1100), Ubuntu, setup's ROCm 7.10 wheel: ~0.8 tok/s decode, GPU at 100% / ~320 W
open · @ibrokers77882 · 3 commentaires · Sur GitHub
BenchmarksSetup & installServer & APIAMD / HIPModels & quantsDocumentation
Description
**AMD RX 7900 XTX (gfx1100): decode ~0.76 tok/s, prefill ~1 tok/s; every GPU stage uniformly slow, raw GPU bandwidth healthy** **System:** Ubuntu [Ubuntu 26.04 LTS], kernel [7.0.0-34-generic], Strata commit [1678de333d0e0711bc414ad992b640e1a37dd814] (engine 0.1.34). Ryzen 9 7900X, 96 GB RAM, RX 7900 XTX 24 GB (gfx1100), display on the iGPU (XTX has ~0 desktop VRAM use). setup.py HIP build with the TheRock 7.10.0a20251120 gfx110X-dgpu wheels, clang 22, libstdc++-16-dev. PCIe Gen4 x16. **Config:** IQ3_S, `--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --kv int8 --kv-resident 32768 --max-context 131072 --pcie-frac 0 --mmap-experts --resident-cpu-experts`. **Symptom:** 58-token prompt takes ~50 s to prefill (~1 tok/s); decode 0.76 tok/s (2.5 tokens/window, ~3300 ms/window). Docs list ~55 tok/s for this card, i.e. ~70x slower. GPU shows 100% busy at 240-390 W while generating; 0% / 20 W idle. **STRATA_DECODE_TIMING:** 3300.82 ms/window = verify 3045 (GPU-reach wait 2695, per-layer host 21 ms) + draft 253. CPU experts 6.76, VRAM hits 21.46, PCIe 0.00 per layer-window. waitCPU 0.27 ms. **STRATA_VERIFY_PROFILE (ms/window, GDN layers):** q8+qkv gemv 489, z 278, out-proj 234, hc-read1+router 323, shared+quant 185, VRAM hits 516, head 368. Nothing dominates; all stages are slow by a similar factor. **Ruled out:** - GPU/driver/clocks: standalone HIP copy test on the same card: 687 GB/s (hipMemcpy D2D), 676 GB/s (own kernel, built with the same clang 22). mclk at max state, performance level auto, no resets or timeouts in dmesg. - Memory mode: same speed in the default pinned-arena mode and with `--mmap-experts --resident-cpu-experts`. - PCIe (stage ~1%, link Gen4 x16), swap and disk I/O, CPU threads, expert hit rate (~85%), VRAM spill (same speed with 2.9 GiB free), GTT limit. - iGPU: HIP_VISIBLE_DEVICES=0, iGPU idle (0%) during generation. - User env: HSA_OVERRIDE_GFX_VERSION=11.0.0 and ROCBLAS_USE_HIPBLASLT=1 were set in my shell; removing both (verified absent in the engine's /proc environ) did not change timing (1m49s for 40 tokens either way). - Build type: Release. Libraries: engine loads only the wheel's libamdhip64/hsa/comgr/rocblas/hipblaslt. **Could not profile:** `rocprofv3 --attach` fails (`couldn't dlsym rocprofiler_register_attach`); running the server under `rocprofv3 --kernel-trace` segfaults in `rocprofiler::hsa::WriteInterceptor` during hipGraphLaunch. **Questions:** Which ROCm/clang versions did you measure on gfx1100 (docs mention ROCm 7.1.52802 / clang 20)? Is a newer wheel index for gfx110X known to work? Is there a built-in kernel benchmark I can run to find which kernel is slow? [strata-log-clean.txt](https://github.com/user-attachments/files/32993476/strata-log-clean.txt) [redacted-stacks1.txt](https://github.com/user-attachments/files/32993541/redacted-stacks1.txt) [redacted-stacks3.txt](https://github.com/user-attachments/files/32993551/redacted-stacks3.txt)
Sur le site
Liens install, modèles, releases.