Issues / #590

#590 RX 7900 XTX (gfx1100), Ubuntu, setup's ROCm 7.10 wheel: ~0.8 tok/s decode, GPU at 100% / ~320 W

open · @ibrokers77882 · 3 comentários · No GitHub

BenchmarksSetup & installServer & APIAMD / HIPModels & quantsDocumentation

Descrição

**AMD RX 7900 XTX (gfx1100): decode ~0.76 tok/s, prefill ~1 tok/s; every GPU stage uniformly slow, raw GPU bandwidth healthy**

**System:** Ubuntu [Ubuntu 26.04 LTS], kernel [7.0.0-34-generic], Strata commit [1678de333d0e0711bc414ad992b640e1a37dd814] (engine 0.1.34). Ryzen 9 7900X, 96 GB RAM,
RX 7900 XTX 24 GB (gfx1100), display on the iGPU (XTX has ~0 desktop VRAM use). setup.py HIP build with the
TheRock 7.10.0a20251120 gfx110X-dgpu wheels, clang 22, libstdc++-16-dev. PCIe Gen4 x16.

**Config:** IQ3_S, `--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --kv int8 --kv-resident 32768
--max-context 131072 --pcie-frac 0 --mmap-experts --resident-cpu-experts`.

**Symptom:** 58-token prompt takes ~50 s to prefill (~1 tok/s); decode 0.76 tok/s (2.5 tokens/window,
~3300 ms/window). Docs list ~55 tok/s for this card, i.e. ~70x slower. GPU shows 100% busy at 240-390 W
while generating; 0% / 20 W idle.

**STRATA_DECODE_TIMING:** 3300.82 ms/window = verify 3045 (GPU-reach wait 2695, per-layer host 21 ms) + draft 253.
CPU experts 6.76, VRAM hits 21.46, PCIe 0.00 per layer-window. waitCPU 0.27 ms.

**STRATA_VERIFY_PROFILE (ms/window, GDN layers):** q8+qkv gemv 489, z 278, out-proj 234, hc-read1+router 323,
shared+quant 185, VRAM hits 516, head 368. Nothing dominates; all stages are slow by a similar factor.

**Ruled out:**
- GPU/driver/clocks: standalone HIP copy test on the same card: 687 GB/s (hipMemcpy D2D), 676 GB/s (own kernel,
  built with the same clang 22). mclk at max state, performance level auto, no resets or timeouts in dmesg.
- Memory mode: same speed in the default pinned-arena mode and with `--mmap-experts --resident-cpu-experts`.
- PCIe (stage ~1%, link Gen4 x16), swap and disk I/O, CPU threads, expert hit rate (~85%), VRAM spill
  (same speed with 2.9 GiB free), GTT limit.
- iGPU: HIP_VISIBLE_DEVICES=0, iGPU idle (0%) during generation.
- User env: HSA_OVERRIDE_GFX_VERSION=11.0.0 and ROCBLAS_USE_HIPBLASLT=1 were set in my shell; removing both
  (verified absent in the engine's /proc environ) did not change timing (1m49s for 40 tokens either way).
- Build type: Release. Libraries: engine loads only the wheel's libamdhip64/hsa/comgr/rocblas/hipblaslt.

**Could not profile:** `rocprofv3 --attach` fails (`couldn't dlsym rocprofiler_register_attach`); running the server
under `rocprofv3 --kernel-trace` segfaults in `rocprofiler::hsa::WriteInterceptor` during hipGraphLaunch.

**Questions:** Which ROCm/clang versions did you measure on gfx1100 (docs mention ROCm 7.1.52802 / clang 20)?
Is a newer wheel index for gfx110X known to work? Is there a built-in kernel benchmark I can run to find which kernel is slow?

[strata-log-clean.txt](https://github.com/user-attachments/files/32993476/strata-log-clean.txt)

[redacted-stacks1.txt](https://github.com/user-attachments/files/32993541/redacted-stacks1.txt)

[redacted-stacks3.txt](https://github.com/user-attachments/files/32993551/redacted-stacks3.txt)

No site

Links install, modelos, releases.