Pull requests / #1320
#1320 test: qualify existing exact HIP paths on gfx906
open · @0FL01 · 0 コメント · GitHub で見る
BenchmarksSetup & installAMD / HIPModels & quants
本文
## Summary Adds regression coverage for existing fast paths. Runtime kernels and defaults are unchanged. - Compare active Q8_1 bytes as well as expert outputs in native_grouped_parity; keep its existing benchmark and output-coverage guard. - Add the real Q2_0 2560x640 shape and a fused gfx906 CTest. - Add multi-row BF16 GEMV checks against independent single-row calls: 1–8 tokens, nine shapes, padded strides, three finite magnitude ranges, graph capture and auxiliary rows. - Enable the existing GDN recurrence test for the separate gfx906 backend. ## Validation Fresh validation checkout based on main 82f46a8, test commit 8f44bf2. Four targeted CTests passed on each of two gfx906 16 GiB cards (8/8). Production stayed online. Build prerequisite: the validation checkout carried the existing gfx906 compatibility fixes in strata_hip.h, vmm.cpp and fused_gr.cu; those are excluded here (related: #1083). No #1187 kernel changes were used in this test build. Full CTest suite and NVIDIA hardware were not tested. The GDN contract checks consumed FP16 outputs and recurrent state; its discarded FP32 scratch can differ. ## Context This does not duplicate #1187 or enable a global performance preset. The existing MMVF_ROWS switch measured +4.07% TG at 4K and +3.13% at 64K on our Hybrid setup, with identical output IDs. Those performance measurements used the #1187 build, not pristine main; [full evidence and limitations](https://github.com/0FL01/Strata/tree/82f4473fb56fd8cf7355814c9512e21d67f4e0d2/bench/results/2026-10-07-gfx906-campaign). Developed with an AI coding assistant; tests and measurements ran on the actual hardware.
関連リンク
インストール・モデル・リリースへの站内リンク。