Pull requests / #744

#744 HIP: on-device autotuning for gfx1100 decode (kernel shapes, draft and CPU settings)

closed · @xyzzing · 0 commentaires · Sur GitHub

BenchmarksAMD / HIPModels & quantsSecurityWindows

Description

## Summary

On-device decode autotuning for the AMD HIP backend (gfx1100): per-shape rows-per-block tuning for the decode expert kernels and `iq_mmvq` (bitwise-exact variants only), a guarded tuning-table format, an on-device tuner, and an orchestrator that sweeps kernels, draft floor, CPU workers, and drafting windows with the user's own prompts. Settings-layer sweeps (`--spec`, `--suffix-draft`, `--spec-min-p`, workers) now work on HIP too, and `calibrate` is HIP-aware (keeps the config's `--pcie-frac`).

Both layers keep answers exactly the same: kernel variants are byte-compared on-device and end-to-end; the settings layer is greedy-verification-exact.

## Measured on the target hardware (RX 7900 XTX 24 GB, IQ3_S, ROCm 7.1.1)

- **Kernel layer: the shipped defaults already win on every IQ3_S shape** (deltas ≤1%, below the keep threshold) — an honest null; the sweep took ~1 min + 4 engine starts.
- **Settings layer: defaults fastest** (57.0 tok/s on the tool's workload).
- **Draft floor: `--spec-min-p 0.70` = 63.1 tok/s vs baseline (~+10%), token-identity verified end-to-end** — adopted in our serving config.
- 128K-context one-shot and ROCm-nightly comparisons are in progress on the same box; the tool's report (`strata-<model>.autotune.json`) carries every measurement.

## Notes

- Base is this fork's HIP integration main (`4da9c59`); a rebase onto current upstream main is queued (the files the patch touches are untouched by 0.1.36–0.1.38's kernel changes, so the rebase should be mechanical).
- The free-VRAM pre-check over-counts (section 6.1 of the authoring notes) — known, fix queued.
- gfx1100 measured; the table's arch gate keeps other cards on defaults until tuned.

Sur le site

Liens install, modèles, releases.