Pull requests / #1360

#1360 gfx1151: STRATA_EXPERT_V2K on by default (UD-Q4_K_XL decode experts: same text, +6.5% greedy / +9.9% sampled output on Strix Halo)

open · @routhjim · 0 评论 · 在 GitHub 查看

BenchmarksSetup & installServer & APIAMD / HIPModels & quantsDocumentationWindowsLinux

描述

`STRATA_EXPERT_V2K` (04e1e11a) runs UD-Q4_K_XL's decode experts on the S26 grouped structure: Q4_K / Q5_K gate/up with Q5_1 / Q8_0 down, 4 rows per warp, the chunk's q8_1 rows in LDS, fused SwiGLU + q8_1. By its design every output is bitwise equal to the native kernels' (`S27<TY>::load/apply` split from the same vec_dot functions). It is opt-in today, and nothing published has measured it.

On gfx1151 it is a clear win, so this PR adds it to the gfx1151 default table (`src/core/arch_defaults.cpp`) next to `STRATA_EXPERT_V2`, and to the section-4 table of `docs/STRIX_HALO.md`. As with every entry there, a user's own setting wins: `STRATA_EXPERT_V2K=0` keeps the old kernels, and `STRATA_GFX1151_DEFAULTS=0` turns the whole table off. The switch only matters for packs whose experts are those formats (UD-Q4_K_XL); other packs never reach it.

**Measured** on a Ryzen AI Max+ 395 (Radeon 8060S, gfx1151, 128 GB, Linux 7.2, ROCm 7.14.1, the iGPU alone).
- Setup: 82f46a8 plus this PR, built per `docs/STRIX_HALO.md` with the section-4 defaults and `STRATA_PF_FUSED_KQ=1`, Unsloth UD-Q4_K_XL, `--mmap-experts --expert-cache auto --spec 4 --mtp mtp/rt --lookup-chain 3 --kv int8`.
- The same binary twice, with `STRATA_EXPERT_V2K=0` for the old kernels.
- Requests: a 1.3K-token prompt with a fresh prefix each time, 512 output tokens, thinking off.
- Timings from `STRATA_DECODE_TIMING=1`; ± is the standard error over the requests.

| | `STRATA_EXPERT_V2K=0` | default with this PR | |
|---|---:|---:|---:|
| output tok/s, greedy (6 requests) | 49.6 ± 0.3 | **52.8 ± 0.8** | **+6.5%** |
| output tok/s, temperature 1.0 / top_p 0.95 / top_k 20 (12 requests) | 41.3 ± 0.5 | **45.4 ± 0.5** | **+9.9%** |
| verify per window (ms) | 54.6 | **51.6** | -3.0 |
| draft per window (ms) | 8.76 | 8.75 | |
| drafts accepted (sampled / greedy) | 52.3% / 66.7% | 51.9% / 65.9% | unchanged |

The draft time and the acceptance do not move: the gain is the verify window's expert step. In a per-stage profile (`STRATA_VERIFY_PROFILE=1`, a Q4_K-dense requant of the same experts) the routed experts went from 25.2 to 21.0 ms per window.

**Same text:** three fixed prompts, greedy, 256 tokens, each twice, give the same text with and without the switch, and across the repeats.

**Not tested:**
- Windows, and other gfx11 parts (the table is gfx1151-only).
- Longer contexts: these runs used a 1.3K-token prompt. The expert step does not depend on the context length, but I did not measure it there.

Measured with the help of Claude (Anthropic).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。