Pull requests / #546
#546 prefill: the fused layout's smaller MoE buffers only when every layer takes the fused path (STRATA_PF_FUSED=1 on IQ2_XS wrote past them)
closed · @sergqwer · 0 コメント · GitHub で見る
AMD / HIPNVIDIA / CUDAModels & quants
本文
Under `STRATA_PF_FUSED=1`, the shipped **IQ2_XS** pack gives a wrong answer after a long prompt. The prompt path writes past its MoE buffers. **Cause** (`src/prefill/prefill.cpp`): - `fused_layout()` switches a streamed chunk (T >= `stream_all_min()`) to the fused path's smaller MoE buffers when **any** layer can run the fused kernels. That is the `fused_ring()` test. - In `moe_bufs(..., fused=true)`, GU / H / Xq / Hq then hold MMQ rows for only `stream_all_min() - 1` tokens. - The path is chosen per layer, though (`fused_l` / `fused_nat`). A layer the native kernels do not cover runs MMQ or the FP16 path over the chunk's T x K rows, and runs on into the next buffers. - The IQ2_XS pack is such a pack. Its 45 IQ2_XXS / IQ2_S layers are covered; its three IQ1_M layers are not, and take the FP16 path. Nothing is reported, and the MoE output is silently corrupted. **Fix:** the smaller layout only when every layer takes the fused path: each layer is in the MMQ plan and, for a native pack, `fused::native_supported()`. - A mixed pack keeps MMQ-sized buffers. Its covered layers still run the fused kernels in them: the log still says "prompt experts on the fused int8 kernels". - A pack whose every layer is covered, and the Q2_0 pack, keep the smaller layout. Nothing changes for them. - Without `STRATA_PF_FUSED=1` nothing changes at all. **Measured** on an RTX 5090: IQ2_XS (ISTA), 262K, `--expert-cache 14900`, first-token logits. The reference is the same engine without the flag (MMQ): | prompt | v0.1.37, `STRATA_PF_FUSED=1` | this PR, `STRATA_PF_FUSED=1` | | --- | --- | --- | | 2K (1 chunk) | KL 0.0039, max abs diff 2.23 | KL 0.0014, max abs diff 0.36 | | 32K (4 chunks of 8192) | **KL 3.79, a different top token**, max abs diff 9.56 | KL 0.0036, same top token, max abs diff 0.69 | | prompt time, 32K | 5,676 ms | 5,134 ms | The remaining KL is the fused kernels' own rounding against MMQ, which the release notes put at MMQ's level. #519 runs `STRATA_PF_FUSED=1` on an IQ3_XXS pack in production: if that pack mixes formats by layer, its long prompts were affected too. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
関連リンク
インストール・モデル・リリースへの站内リンク。