Pull requests / #1661
#1661 Perf/gfx906 qsa query swizzle
open · @0FL01 · 0 commentaires · Sur GitHub
BenchmarksMulti-GPUAMD / HIPModels & quantsSecurityDocumentation
Description
## Summary Add a default-off query-LDS layout permutation for INT8 batched QSA attention on gfx906 wave64. Enable with `STRATA_GFX906_ATTN_QUERY_SWIZZLE=1`. The query store uses `d ^ ((d >> 3) & 4)` and the two vector loads recover the original logical operands. The 12 KiB query tile, 16-byte alignment, floating-point expressions and operation order are preserved. Score reduction, softmax, value accumulation and merge code are unchanged. The runtime guard checks the calling thread's current device, accepts exact gfx906 architecture names with valid feature suffixes, requires wave64, and fails closed on HIP errors. Property caching is per thread and last device ordinal. ## Scope - Applies to INT8 `qsa_decode_attn_batch` calls in PREFILL, VERIFY and MTP, including `n_q=1`. - A k8v4 main model can reach the INT8 path through its MTP drafter. - Keeps `qsa_decode_attn_step`, other pool formats, lane-cell scoring and other GPU/default paths unchanged. - Adds a deterministic component probe and a source-extracted host-guard test. ## Validation - 31 host-only mock-HIP guard checks passed, covering opt-in and format gating, architecture/wave rejection, device switching, thread-local caching, API failures and recovery. - Both gfx906 GPUs passed raw-byte parity for `n_q=1/context=4`, `n_q=33/context=4096` with masked pages, and `n_q=32/context=65536`. Two flag states × three geometries × two GPUs × two builds (hardened experimental and clean upstream) give 24 runs, or 12 on/off comparisons. - Final clean-upstream A/B/B/A component medians at `n_q=32/context=65536`, 50 repetitions per run: A = 1.188960/1.187439 ms; B = 1.036240/1.036560 ms. Throughput gain from the ratio of arm means is **14.646806%**, including chunk and merge. - Compiled kernels retain 24 `ds_read_b128` query loads, 15,872 LDS bytes, 42 SGPRs, four waves per SIMD and zero scratch/spills. VGPRs increase from 41 to 42; encoded static instructions increase from 1,000 to 1,005. - Guard hardening preserves all disassembled device text and AMDGPU resource notes relative to the previously qualified experimental kernel. The initial component result of 14.791570% missed the predeclared 20% gate. An explicit bounded cost-benefit exception allowed one full-model 65,536-input/1,024-output comparison. With the same experimental HC + grouped-MMQ + PR #1525 dequantization baseline in both arms, toggling only QSA yielded **+1.185195% prompt throughput**, **−1.171313% native prompt time** and **−1.048564% request wall time**; all 1,024 output IDs matched. The +0.418279% generation movement is noise-scale. This single pair predates host-guard hardening and was not run on rebased upstream main; the device-code audit above covers the hardening step. The final hardened combined engine (`3d2650fe62fdea24679ac90401eac505b541ccc47886e59885a639c438cb4c17`) subsequently completed a 200,000-token prompt plus 256 generated tokens at capacity 204,800, matching all archived stable-control output IDs. This was a capacity/parity test, not a fresh speed comparison. With the actual executable hash and candidate flags verified, targeted API checks passed for a synthetic forced tool/result roundtrip, two CPU-vision image requests and text afterward, unauthenticated 401/root 404 responses, and A/auxiliary/A RAM cache restoration. Warm and restored requests reused 6,423 tokens. Original configuration and Compose files were restored byte-for-byte, with the stable API healthy afterward. Full measurements, test receipts and build provenance: [benchmark report](https://github.com/0FL01/Strata/blob/1f555de86254cc01ec06882c70f9584c1f09ec39/bench/results/2026-10-08-gfx906-qsa-query-swizzle/README.md). ## Reproduction scope and limitations Clean base: `fb58e0dbc8399662c0e47c76578c6e878b14f6cf`. Final clean patch SHA-256: `9c81295b8cdf8354fe0e5bc6de36a547199becd1a8fc2a79daabf2b9da9548c6`. Tests use two gfx906 wave64 GPUs (runtime reports AMD Radeon Pro VII), llama.cpp `3cf03257f219afbe7334045ff7c6a06ac68c627d`, and the pinned ROCm build image `sha256:bccb7ee7e7a78274519db9a43ba63c34ddd2e74bb60f8764a8f50aaee1f2c646`. The compiler is patched AMD Clang `23.0.0git`; the ROCm distribution version is unrecorded. The clean Release probe was built with portable/native-expert settings and MMQ build options disabled. Performance and numerical evidence are limited to the tested inputs and build stack. Run the host guard with `python3 bench/test_gfx906_qsa_query_swizzle_guard.py`. Build the `gfx906_qsa_query_swizzle_probe` CMake target and compare flag-off/on outputs in separate processes; the dispatcher caches the environment setting. The timing probe invocation is `gfx906_qsa_query_swizzle_probe 32 65536 50 output.bin`. Keep lane-cell scoring disabled and compare full output files. The capacity and API/cache tests qualify the frozen experimental stack. The clean rebased patch has component and host-guard validation only; no upstream CI pass, full-model run on rebased main or production promotion is claimed. ## Integration note PR #1402 introduces a third `attn_chunk_kernel` template argument `G`; this patch uses a third boolean argument. These changes need coordination before merging. No direct duplicate of the query-LDS swizzle was found in the overlap review.
Sur le site
Liens install, modèles, releases.