Pull requests / #1535

#1535 gfx906: add guarded register top-k for decode

open · @0FL01 · 0 コメント · GitHub で見る

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows

本文

## Summary

Add an opt-in gfx906 decode route using the existing register-66 top-k kernel at larger configured capacities. Enable it with `STRATA_GFX906_TOPK_REG=2`; defaults stay unchanged.

The route reads current device steps during execution, so captured graphs can cross the context threshold and return to a shorter restored conversation.

## What changed

- Limit the extension to gfx906, uncounted decode, and windows of at most eight queries.
- Use complementary device-side guards: the final query at or below 24,576 cells selects the reference kernel; longer contexts select register66. Exactly one arm writes the whole window, with guards before barriers.
- Preserve the existing smaller-capacity route, oversized-capacity fallback, `STRATA_TOPK_OLD` override, and counted/prompt behavior.
- Add a guarded CTest, adversarial score inputs, output sentinels, and changed-step graph replay across threshold crossings and restores.

Only three implementation/test/CMake files change. Q2 GPU/CPU work and the broader parity changes from #1187/#1320 are excluded.

## Validation

Built the full engine and parity target on upstream `6674a006` with the same pinned gfx906 HIP image and ggml revision for both arms.

On both 16-GiB gfx906 cards:
- Both registered top-k CTests pass.
- 60 driver cases pass in total, including the two CTest invocations.
- Coverage includes query windows 1/6/8/9, exact threshold neighbors, contexts up to 204,800 in graph replay, counted calls, the old override, and an oversized-capacity fallback.

Isolated fresh-process model ABBA, 1,024 generated tokens/run, identical settings and output IDs within every workload:
- Code 4K: 53.525 → 53.428 tok/s (**−0.18%**)
- Code 64K: 47.730 → 48.295 tok/s (**+1.18%**)
- Russian 64K: 49.747 → 50.447 tok/s (**+1.41%**)

Prefill is effectively unchanged. These are two observations per arm on this hardware/model, not a broad performance guarantee. The small short-context cost is one reason to keep this opt-in.

[Full results, test arguments, logs, build provenance and limits](https://github.com/0FL01/Strata/blob/ca3533157d066a13b128f3daa61f56411dcf3155/bench/results/2026-10-08-gfx906-guarded-topk/README.md).

## Extra notes

The base's full gfx906 build needs the one-line `cudaEventBlockingSync` alias already proposed in #1396. That same build-only prerequisite was applied to **both** measured arms and is not duplicated in this PR. Its separate Q6_K change was not needed.

CUDA and non-gfx906 HIP builds/runtime were not run here. Static portability review changed the test's capture mode to the ThreadLocal constant already mapped by both HIP compatibility headers. No actual 200K model request is claimed for this isolated A/B; the component graph tests cover that context.

Draft for upstream review.

関連リンク

インストール・モデル・リリースへの站内リンク。