Issues / #1549
#1549 Prompt reading is 72x slower when experts come from the host mirror, and ~48x off the PCIe bound
open · @demetree · 3 comments · View on GitHub
BenchmarksMulti-GPUModels & quantsDocumentationWindows
Description
### The prompt path collapses when experts come from the host mirror, and not at the PCIe bound The host mirror (plan item 2: experts that do not fit in VRAM read from pinned host memory over PCIe instead of the SSD) works well for decode. It does not work for the prompt, and the gap is much larger than bandwidth can explain. Arc Pro B70 32 GB, Windows, OpenCL backend, Qwen3.8-Flash-Next-GSQ-RCO, `--spec 4 --mtp`, greedy, medians of two interleaved runs where both rows are the same build: | | experts in VRAM | decode, 2,048-token prompt | decode, 3-token prompt | prompt reading, 2,048 tokens | |---|---|---|---|---| | IQ1_M | 23.42 GiB, **all resident** | 43.6 tok/s | 76.4 tok/s | **74.7 tok/s** | | IQ3_XXS | 39.97 GiB, 17,687 of 24,576 resident + 6,889 (11.15 GiB) mirrored | 30.0 tok/s | 19.2 tok/s | **1.03 tok/s** | Decode absorbs the misses - 30 tok/s with 28% of the experts over PCIe is a reasonable price, and the device-built verify plan points the expert kernels straight at the mirror. The prompt does not: 1.03 tok/s is a 72x drop, and one 2,048-token prompt took 1,987,881 ms (33 minutes), of which 1,988,700 ms was the wait for the first token - it is all in the prefill, which the 4,096-row chunk handles in one pass. The ratio is the interesting part. If each layer simply streamed the mirrored experts it needed over PCIe, 11.15 GiB per layer across 48 layers is ~535 GiB, which at the 13.3 GB/s host-to-device I measured on this machine is ~41 s, not 1,988 s. So the mirror's prompt reads are ~48x off the bandwidth bound: they do not look pipelined the way the decode plan's are. For reference, `Prefill::prefill.cpp` already carries a note that "part of the model in the host mirror a short first chunk costs a whole extra pass of that stream", and `g_pinned_share` drops the staging ring to 96 slots when less than 90% of the streamed experts are pinned - the design expects streamed experts to go through a small host-copy ring, so a *fully pinned* mirror may be hitting a path that was tuned for the other case. ### What I checked The mirror is wired into both paths, so this is not a missing wire: `resident_plan_set_mirror()` gives the verify plan its per-layer mirror addresses, and the prefill reads a mirrored expert with `m.src->pinned(l, e)` -> `m.src->blob(l, e)` rather than staging it. The engine also *knows* the difference is not bandwidth-sized - the existing all-resident forced test in docs/INTEL.md measures 3.6 GiB of misses costing ~6% of decode, and 3.6 GiB here would be a few percent, not 72x. ### What would help A profile of one prefill chunk with and without the mirror would say where the 48x goes - most likely either the per-expert launches not being enqueued far enough ahead, or the chunk re-reading the same mirrored expert once per row-block. The number that would confirm it is bytes moved per second over PCIe during the prefill: 535 GiB in 1,988 s is 270 MB/s, about 2% of the link. For anyone choosing a model on a card this size in the meantime: a model that fits entirely in VRAM is fine on both paths, and one that does not is decode-usable and prompt-useless. That is worth knowing before choosing one.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.