Pull requests / #1187

#1187 gfx906: opt-in Q2_0 gate/up weight reuse without singleton regression

open · @0FL01 · 0 コメント · GitHub で見る

Server & APIModels & quants

本文

## Change
- Adds opt-in `STRATA_EXP_MODE=13` for `Q2_0 gate/up` on `gfx906`, `H=2560 / FF=640`.
- Reuses weights across up to four tokens without LDS staging, retaining the baseline 16-row block geometry.
- Default mode `7` and `down projection` remain unchanged.

## Testing
- Fresh-engine ABBA on two `16 GiB gfx906 GPUs`: `TG +1.48%` at `4K` and `+1.43%` at `64K`, with all `1024` output IDs identical in each code workload.
- `280` component cases and `282` stronger intermediate-parity cases passed.
- No measured singleton regression in the tested grid.

## Notes
- Russian baseline runs were themselves nondeterministic, so that timing is not equal-work evidence.
- Existing `gfx906` build fixes (#1083) are excluded.
- Related PRs were reviewed; no direct duplicate found.
- Developed with an OpenAI coding assistant.

関連リンク

インストール・モデル・リリースへの站内リンク。