Pull requests / #744
#744 HIP: on-device autotuning for gfx1100 decode (kernel shapes, draft and CPU settings)
closed · @xyzzing · 0 comentarios · En GitHub
BenchmarksAMD / HIPModels & quantsSecurityWindows
Descripción
## Summary On-device decode autotuning for the AMD HIP backend (gfx1100): per-shape rows-per-block tuning for the decode expert kernels and `iq_mmvq` (bitwise-exact variants only), a guarded tuning-table format, an on-device tuner, and an orchestrator that sweeps kernels, draft floor, CPU workers, and drafting windows with the user's own prompts. Settings-layer sweeps (`--spec`, `--suffix-draft`, `--spec-min-p`, workers) now work on HIP too, and `calibrate` is HIP-aware (keeps the config's `--pcie-frac`). Both layers keep answers exactly the same: kernel variants are byte-compared on-device and end-to-end; the settings layer is greedy-verification-exact. ## Measured on the target hardware (RX 7900 XTX 24 GB, IQ3_S, ROCm 7.1.1) - **Kernel layer: the shipped defaults already win on every IQ3_S shape** (deltas ≤1%, below the keep threshold) — an honest null; the sweep took ~1 min + 4 engine starts. - **Settings layer: defaults fastest** (57.0 tok/s on the tool's workload). - **Draft floor: `--spec-min-p 0.70` = 63.1 tok/s vs baseline (~+10%), token-identity verified end-to-end** — adopted in our serving config. - 128K-context one-shot and ROCm-nightly comparisons are in progress on the same box; the tool's report (`strata-<model>.autotune.json`) carries every measurement. ## Notes - Base is this fork's HIP integration main (`4da9c59`); a rebase onto current upstream main is queued (the files the patch touches are untouched by 0.1.36–0.1.38's kernel changes, so the rebase should be mechanical). - The free-VRAM pre-check over-counts (section 6.1 of the authoring notes) — known, fix queued. - gfx1100 measured; the table's arch gate keeps other cards on defaults until tuned.
En el sitio
Enlaces a install, modelos, releases.