Pull requests / #546

#546 prefill: the fused layout's smaller MoE buffers only when every layer takes the fused path (STRATA_PF_FUSED=1 on IQ2_XS wrote past them)

closed · @sergqwer · 0 comments · View on GitHub

AMD / HIPNVIDIA / CUDAModels & quants

Description

Under `STRATA_PF_FUSED=1`, the shipped **IQ2_XS** pack gives a wrong answer after a long prompt. The prompt path writes past its MoE buffers.

**Cause** (`src/prefill/prefill.cpp`):
- `fused_layout()` switches a streamed chunk (T >= `stream_all_min()`) to the fused path's smaller MoE buffers when **any** layer can run the fused kernels. That is the `fused_ring()` test.
- In `moe_bufs(..., fused=true)`, GU / H / Xq / Hq then hold MMQ rows for only `stream_all_min() - 1` tokens.
- The path is chosen per layer, though (`fused_l` / `fused_nat`). A layer the native kernels do not cover runs MMQ or the FP16 path over the chunk's T x K rows, and runs on into the next buffers.
- The IQ2_XS pack is such a pack. Its 45 IQ2_XXS / IQ2_S layers are covered; its three IQ1_M layers are not, and take the FP16 path. Nothing is reported, and the MoE output is silently corrupted.

**Fix:** the smaller layout only when every layer takes the fused path: each layer is in the MMQ plan and, for a native pack, `fused::native_supported()`.
- A mixed pack keeps MMQ-sized buffers. Its covered layers still run the fused kernels in them: the log still says "prompt experts on the fused int8 kernels".
- A pack whose every layer is covered, and the Q2_0 pack, keep the smaller layout. Nothing changes for them.
- Without `STRATA_PF_FUSED=1` nothing changes at all.

**Measured** on an RTX 5090: IQ2_XS (ISTA), 262K, `--expert-cache 14900`, first-token logits. The reference is the same engine without the flag (MMQ):

| prompt | v0.1.37, `STRATA_PF_FUSED=1` | this PR, `STRATA_PF_FUSED=1` |
| --- | --- | --- |
| 2K (1 chunk) | KL 0.0039, max abs diff 2.23 | KL 0.0014, max abs diff 0.36 |
| 32K (4 chunks of 8192) | **KL 3.79, a different top token**, max abs diff 9.56 | KL 0.0036, same top token, max abs diff 0.69 |
| prompt time, 32K | 5,676 ms | 5,134 ms |

The remaining KL is the fused kernels' own rounding against MMQ, which the release notes put at MMQ's level. #519 runs `STRATA_PF_FUSED=1` on an IQ3_XXS pack in production: if that pack mixes formats by layer, its long prompts were affected too.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.