Pull requests / #1042

#1042 prefill: bytes_needed sizes the MoE buffers for the layout init picks (with or without an expert source)

closed · @sergqwer · 0 コメント · GitHub で見る

NVIDIA / CUDA

本文

Replaces #547. GitHub closed it on 2026-10-05, when my fork was made private by mistake: that took the fork out of the network for good, so it can no longer open pull requests. This is the same branch and commit (`b3a15cf`), opened from a new fork; the discussion and the measurements are in #547.

---

`Prefill::bytes_needed()` sizes the region the prompt path borrows from the expert cache, and it must count exactly what `carve()` allocates. It does not, in two places (`src/prefill/prefill.cpp`).

1. **`xn` and `grs`.** `bytes_needed` counts a T x D float buffer for `xn`. `carve` allocates it only under `STRATA_GR_UNFUSED=1`; by default the fused read reads `R`. `bytes_needed` also leaves out `grs` (T x HC floats), which `carve` always takes. Net, 40,944 bytes a token too many:
   - 0.31 GiB at the default 8192-token chunk;
   - 1.25 GiB at a 32K chunk (`--prefill auto:32768`).

   Those expert-cache slots are lent to every prompt, never used, and refilled after it. A 12-16 GB card feels it most: either more experts are streamed during the prompt, or the auto chunk steps down earlier than it needs to.
2. **The MoE layout.** `bytes_needed` sizes the MoE scratch with `fused_layout(T, true)`, but `carve` uses `fused_layout(T, m.src != nullptr)`. For a pack that can take the fused layout, a caller with no `ExpertSource` gets a region sized for the smaller fused buffers, and `init` then asks for MMQ-sized ones. `bytes_needed` now takes a `src` flag (default `true`), and `generate.cpp` passes `srcp != nullptr` at its four call sites.

**Measured** on an RTX 5090, IQ2_XS (ISTA), `--expert-cache 14900`, v0.1.37 against this PR:

| prompt | slots the prompt path borrows | first-token logits | prompt time (one run each) |
| --- | --- | --- | --- |
| 2K (1 chunk) | 1,303 -> 1,239 (1.75 -> 1.67 GiB) | byte-identical | 871 / 884 ms |
| 32K (4 x 8192) | 3,340 -> 3,108 (4.49 -> 4.18 GiB) | byte-identical | 6,460 / 6,241 ms |

The same experts are computed with the same kernels; only the number of slots lent for the prompt changes.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

関連リンク

インストール・モデル・リリースへの站内リンク。