Pull requests / #1297
#1297 HIP: on-device autotuning for gfx1100 decode (kernel shapes, draft and CPU settings)
open · @xyzzing · 0 comentários · No GitHub
BenchmarksSetup & installAMD / HIPModels & quantsSecurityDocumentationWindows
Descrição
Re-file of #744 — that PR was auto-closed on 2026-10-06 when this repository's main history was cleaned up (per the maintainer's note there, not a rejection). Original discussion: #744. One deliberate change from the original PR: the original branch carried the 0.1.27-era patch (`6947459`) as-is. This re-file ships the **ported and measured version instead** (fork-side, 2026-10-05): the tool re-applied onto the current engine line, with the kernel side re-checked against what upstream evolved since 0.1.27 (the recorded on-card campaign kept every shipped kernel shape — the tuner's value on this machine was proving it per-machine, not changing anything). Docs (`docs/AMD_HIP_AUTOTUNE.md` + the `docs/AMD_HIP.md` pointer) are included; the fork-only doc row from the original branch is not. Re-base notes, explicitly: `tools/calibrate.py` grew `host_worker_extras()`/`extra_workers` on main since the port; the reconciliation keeps main's CPU-worker machinery and adds the port's backend/draft/sweep surface on top (the stub tests updated to match). The (fork-only) GFX1100.md row of the original branch is dropped as noted. Gates on this base: full HIP build green; `decode_tuning_test` all checks pass; `tools/hip/test_autotune.py` + `tools/test_calibrate.py` 34 passed. The measured verdict from the port campaign stands: on this gfx1100 the sweep kept every shipped kernel-shape default (byte-gated, 7/7) and the settings layer at the shipped defaults (67.1 tok/s vs 65.0/66.1 swept) — the tool's value is proving it per machine. --- ## Summary On-device decode autotuning for the AMD HIP backend (gfx1100): per-shape rows-per-block tuning for the decode expert kernels and `iq_mmvq` (bitwise-exact variants only), a guarded tuning-table format, an on-device tuner, and an orchestrator that sweeps kernels, draft floor, CPU workers, and drafting windows with the user's own prompts. Settings-layer sweeps (`--spec`, `--suffix-draft`, `--spec-min-p`, workers) now work on HIP too, and `calibrate` is HIP-aware (keeps the config's `--pcie-frac`). Both layers keep answers exactly the same: kernel variants are byte-compared on-device and end-to-end; the settings layer is greedy-verification-exact. ## Measured on the target hardware (RX 7900 XTX 24 GB, IQ3_S, ROCm 7.1.1) - **Kernel layer: the shipped defaults already win on every IQ3_S shape** (deltas ≤1%, below the keep threshold) — an honest null; the sweep took ~1 min + 4 engine starts. - **Settings layer: defaults fastest** (57.0 tok/s on the tool's workload). - **Draft floor: `--spec-min-p 0.70` = 63.1 tok/s vs baseline (~+10%), token-identity verified end-to-end** — adopted in our serving config. - 128K-context one-shot and ROCm-nightly comparisons are in progress on the same box; the tool's report (`strata-<model>.autotune.json`) carries every measurement. ## Notes - Base is this fork's HIP integration main (`4da9c59`); a rebase onto current upstream main is queued (the files the patch touches are untouched by 0.1.36–0.1.38's kernel changes, so the rebase should be mechanical). - The free-VRAM pre-check over-counts (section 6.1 of the authoring notes) — known, fix queued. - gfx1100 measured; the table's arch gate keeps other cards on defaults until tuned. *Prepared with an AI engineering agent (GLM-5.3 / Z.ai) under human direction; every number is from our own recorded runs.* ## The second commit: the seam rework 1. **Kernel-selection-space token** **Why:** Between the original PR and this re-file, the decode kernels were reworked for sub-warp/`EXACT_N` behavior while using the same compiler and HIP runtime. The old table identity — arch + runtime + compiler hash — would have accepted a table whose measured choices no longer describe the current build. **What:** The table header gains a 4th token: the kernel-selection-space token. This build reports token `2` (see the wrong-space refusal in the measurements). Tables made before this token existed, or made for a different token, refuse as stale and tell the user to run the tuner again. 2. **Dispatch-knob refusal** **Why:** The tuning table measures the default decode dispatch. If a rerouting knob changes the dispatch path, the measured rows no longer describe the launch being made. **What:** The table refuses to activate when any of these knobs are set: `STRATA_OLD_IQ_MMVQ`, `STRATA_NO_SUB16_GU`, `STRATA_IQ_STAGE_GRID=0`, `STRATA_IQ_STAGE_GRID_MMVQ=1`, `STRATA_EXPERT_V2`/`V2K`, `STRATA_TSUM`, `STRATA_GROUPED_V1`. The refusal is one stderr line; defaults stay in effect. 3. **One rows list** **Why:** The compiled-in rows sets previously lived in two places: parser validity and launcher switches. Those could drift. **What:** Parser, launchers, and the tuner now derive the rows lists from one X-macro pair in `decode_tuning.hpp`: `STRATA_EXPERT_ROWS_VARIANTS` and `STRATA_MMVQ_ROWS_VARIANTS`. 4. **True-default baseline** **Why:** The tuner’s default candidate and bitwise reference were the explicit-rows pair `(8,8)`. That is not the launch the engine actually makes when no table is active. **What:** The default candidate and bitwise reference are now the engine’s real entry points: `native_expert_grouped` / `iq_mmvq` with no table. The default pair also skips the bitwise self-check: comparing it to itself decides nothing, and excluding it would crash the chooser. 5. **Doc updates** **Why:** A table row and the default geometry can both contain a number like `16`, but they do not mean the same mechanism. **What:** Docs now state: - table rows describe plain one-warp-per-row kernels; - the default path uses the multi/sub-warp chain; - `16` in a table and `16` in the default’s geometry are different mechanisms; - the mmvq key uses the token-count blend; - the new stale/space/knob refusals are documented. ## Measurements (RX 7900 XTX, gfx1100, ROCm nightly SDK) | Check | Result | Source | | --- | --- | --- | | `decode_tuning_test` (parser + every refusal, CPU-only) | all checks pass (now includes the stale-3-token-header and wrong-space refusal cases) | `decode_tuning_test` on this branch build `23b447b`; log in footer | | `iq_parity` / `native_grouped_parity` / `iq_multi_parity` on card | 100% passed, each | arm/log `23b447b`; log in footer | | `native_grouped_parity` through an ENGAGED 1-row table | 100% passed; engagement print `decode tuning: 1 shapes from ... (gfx1100, runtime 71726392)` | arm/log `23b447b`; log in footer | | refuse paths on card | wrong space -> `made for kernel selection space 99, this build is 2 (the kernels changed): run the tuner again`; stale pre-space header -> `made before the kernel selection space was tracked - the kernels have changed since`; knob set -> `STRATA_NO_SUB16_GU reroutes the decode dispatch away from what the table measured`; defaults ran in every refusal | arm/log `23b447b`; log in footer | | tune campaign, 7 expert pairs (18:20, 18:42, 21:20, 21:42, 22:20, 22:42, 23:20; n_embd 2560, n_ff 640), default candidate = the engine's own entry | HONEST NULL x7 - `default chain 67.9-73.1 us, best 8/8, kept the default`, 0 shapes changed | arm/log `23b447b`; log in footer | | end-to-end | not run - nothing to install on a null campaign (the e2e gate - tokens identical, decode not slower than 2% - is unchanged from #744) | #744 gate; no new run in this record | arm = `tools/hip/tune_decode` + `autotune.py` on this branch's build (`23b447b`) through the HIP pipeline against the nightly SDK; full log and refusal transcripts in the fork-side record (`artifacts/rocm/results/record-autotune-refile-gpu-20261007.json`). The campaign result is a null: the 0.1.31–0.1.39 kernel evolution and the sub-warp rework's hardcoded conclusions already win here. The tool's value is per-machine verification: it shows the defaults are fastest, byte-checked, and it tunes where they are not. The HIP-runtime field tracks the loaded runtime library: the same binary reported `70152802` against the system `libamdhip64` and `71726392` against the SDK one, reported in the same card record; tune and run with the same library the engine loads.
No site
Links install, modelos, releases.