Issues / #1258

#1258 gfx1151 (Strix Halo) prompt: per-kernel profile of a 16K prompt (rocprofv3) - where would outside help be useful?

open · @huppiflupp · 2 comentários · No GitHub

BenchmarksAMD / HIPModels & quantsDocumentationLinux

Descrição

**Box.** Ryzen AI Max+ 395, 128 GB (gfx1151, Radeon 8060S, 40 CUs, 2.9 GHz), Linux 7.2, ROCm 7.14.1 (TheRock, hipBLASLt 100401), Strata 82f46a8 built per docs/STRIX_HALO.md.
**Run.** UD-IQ4_XS with a Q6_K dense trunk, one 16,058-token prompt, `--prefill 16384 --kv int8 --mmap-experts --expert-cache 24576 --no-prefill-borrow --spec 4 --mtp mtp/rt --mtp-q4 all`, the section-5 switches of STRIX_HALO.md. `rocprofv3 --kernel-trace --stats` (1,193 tok/s under the profiler, 1,230 without).

**GPU time by kernel (12.8 s total, top entries):**

| kernel | calls | ms | share | registers / LDS / scratch |
|---|---:|---:|---:|---|
| `pfg::kernel_bk64` (FP16 prompt projections) | 181 | 2,249 | 17.6 % | 248 VGPR, 54 KB |
| `native_w11_kernel<IQ3_S, gate/up>` | 172 | 1,983 | 15.5 % | 192 VGPR, 31 KB, scratch 12 |
| `prompt_attn_wmma_kernel` | 12 | 1,243 | 9.7 % | 224 VGPR, 22 KB |
| `native_w11_kernel<IQ4_NL, down>` | 172 | 1,023 | 8.0 % | 192 VGPR, 25 KB, scratch 48 |
| `gr_write_norm_rs_kernel` | 94 | 938 | 7.3 % | 32 VGPR |
| `gr_upmix_kernel` | 96 | 866 | 6.8 % | 248 VGPR, 41 KB |
| `gdn_rec_quad2c_kernel` | 36 | 584 | 4.6 % | |
| `moe_combine4_kernel` | 48 | 455 | 3.6 % | |
| `pfg::kernel_hcd_exact` | 96 | 454 | 3.5 % | 248 VGPR, 36 KB |
| `__amd_rocclr_fillBufferAligned` | 477 | 301 | 2.3 % | one call is 297 ms |
| `dequant_kernel<Q6_K>` | 252 | 258 | 2.0 % | |

**Rough efficiency (my arithmetic, please correct):**
- Routed experts: gate/up ~50 TFLOP (16,058 x 10 x 48 x 2 x 2560 x 640 x 2) in 1.98 s = ~25 TFLOP/s; down ~25 TFLOP in 1.02 s = ~25 TFLOP/s, against a WMMA peak of roughly 59 on this part.
- Hyper-connection read/write: `gr_write_norm_rs` moves >= ~1 GB per call (R fp32 in, xn16 out) in 10.0 ms = ~100 GB/s; `gr_upmix` >= ~1 GB per call in 9.0 ms = ~110 GB/s, against ~256 GB/s theoretical. Together with `kernel_hcd_exact` that is 2.26 s (17.6 %), the largest share after the expert kernels.
- The phase timer (STRATA_PREFILL_TIMING) put "gemm down" at 197 ms on the same run type; the trace shows the down kernel at 1,023 ms, so the phase split mixes the expert kernels.

**Question for the maintainers:** you are actively working on gfx1151 (the Aurora port, the W bundle 528c1fb `fused-expert exact wins`). Where would outside help be useful and not get in your way? Candidates I can see from this profile: (a) the memory-bound hyper-connection kernels (norm_rs / upmix at ~40 % of bandwidth), (b) the routed-expert kernels, (c) a unified-memory default that skips the prompt-path loan (see #1236: +5.2 % prompt in the server here). I can run interleaved A/B measurements and bitwise checks on this box for any branch you point me at, and I can test the #1139 fix on gfx1151 with `--parallel`.

The trace CSVs are available if useful.

_Measured and written with an AI assistant (Claude Code)._

No site

Links install, modelos, releases.