Pull requests / #1417
#1417 prefill: scope each split-stage loan estimate to its device
open · @imanu86 · 0 comments · View on GitHub
Multi-GPUAMD / HIPNVIDIA / CUDAWindows
Description
A layer-split prompt borrows buffers from each stage's expert cache. This makes the `part_slots` estimate run under `OnDevice(p.dev)`, matching the device scope used by `Prefill::init` and the other prompt-buffer estimates. When initialization fails inside a borrowed region, the diagnostic also reports that region's capacity. This is a small consistency change, not a demonstrated fix for a failure in current upstream. At upstream `d5ea7133741e67743c0e886bb426c0ce8d69cf6c`, CUDA `prompt_f16()` is always false and the HIP setting is process-wide, so these devices currently produce the same count. The added scope makes the existing contract explicit before device-specific buffers are introduced. The motivating failure occurred in our fork, which has extra Turing-only prompt buffers: RTX 3060 as the first stage and RTX 2080 Ti as the last stage, on Windows. The later stage's loan was undercounted even though free VRAM remained. Those fork-specific buffers are not part of this PR, and their successful run is not validation of this exact upstream branch. Validation and scope: - Branch `pr/split-prompt-borrow`, head `4f9f14892da9e9d401dd7f6360e71b311912858a`, one commit over `82f46a8c8f475f001ad76d92f58f4a4f8ffb0253`. - `git diff --check 82f46a8..HEAD` passes. Read-only inspection against current upstream confirms neither addition is present there. - No fresh build or GPU run of this exact branch was performed for publication. CUDA/HIP compilation and a long-prompt layer-split regression remain pending. - No numerical kernel changes, new option, or speed claim. `OnDevice` restores the previous device when the count ends. Suggested validation is to compare prompt loan sizes and successful prompt ingestion before/after on an ordinary layer split. Reproducing the original undercount requires device-specific allocation differences, which current upstream does not have.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.