Pull requests / #1451

#1451 prefill: allocate PLE host buffers only when needed

open · @W1nge · 0 comentarios · En GitHub

NVIDIA / CUDAModels & quantsWindows

Descripción

Prefill initialization allocates two maximum-chunk PLE host buffers even when the next request does not use batched PLE. Allocate them when a PLE-enabled run starts instead: one for a single chunk, two for multiple chunks, retaining capacity across smaller requests. Before growing a buffer, wait for its previous upload. Keep the existing pageable fallback and forced-pageable path.

At a 32,768-row layout and 2,560 floats per row, each buffer is 320 MiB. This avoids both allocations at initialization and the second allocation for single-chunk runs. Allocation cost moves to the first PLE run; capacity remains allocated after growth. This does not change chunk defaults, numerical kernels, or GPU workspace layouts, and adds no tuning flags.

Validation on Windows, RTX 2080 Ti (sm75), CUDA 12.4, MSVC 19.44:

- 64 real delayed-DMA cases pass, checking downloaded values through growth/reuse and pinned/pageable transitions. The test invokes the production helper directly.
- CMake/Ninja build and CTest pass: `cmake --build build --target ple_host_buffer_test`, then `ctest --test-dir build -R '^ple_host_buffer_test$' --output-on-failure` (`STRATA_ENABLE_CUDA=ON`, `STRATA_BUILD_TESTS=ON`). No model files required; no CUDA device returns skip code 77.
- `prefill.cpp` compiles both without native/MMQ macros and with native/MMQ/fused enabled.

Draft: full-model inference and other GPU backends have not been rerun on this upstream port. Earlier fork tests included other changes, so their throughput gains are not attributed to this isolated patch.


Validation update (2026-10-08): a complete current-main integration build including #1451, #1454, #1455, #1457 and #1465 passed four real IQ3_XXS model baseline/integration pairs on RTX 2080 Ti: default, GR_UNFUSED, PREFILL_BF16X2 and legacy RING_BYTES modes. Each used INT8 KV, 194 input tokens, two prefill chunks and 8 output tokens; all runs exited successfully and output IDs matched within each pair. A 256K configured-context startup and short generation also passed (this did not fill a 256K prompt). These are combined regression checks, not an isolated performance result or exhaustive state equality. Other hardware/backends remain untested. The earlier draft-only status reflected the absence of any current-main model run; now requesting review with these limits explicit.

En el sitio

Enlaces a install, modelos, releases.