Pull requests / #1649
#1649 prefill: take expert blobs from a helper's VRAM cache when it holds them
open · @W1nge · 0 comments · View on GitHub
Setup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindows
Description
A helper GPU's expert cache already holds a copy of the blobs the prompt path streams, and the helpers compute nothing while a prompt is read on a serial request path. On this host a full-size pack keeps about 14 GiB of experts per helper card, so the same bytes are read from the files and then borrowed back over PCIe. Add `RemotePrefillSource`, an `ExpertSource` that wraps the file source and asks the helpers first. A blob a helper holds is copied cache-to-host on that helper's own stream; misses are merged into one call to the fallback's existing batched reads, so the file read stays batched and unchanged. Every other decision — identity, release, the profiled order — is forwarded as it was, and the host copies are the same bytes the file holds, so the prompt path's arithmetic does not change. Opt-in through `STRATA_PREFILL_HELPER_SOURCE`. It applies only where the serve path is serial (one GPU, no layer split, no batch slots) and the prompt path streams every expert the main card does not hold; everywhere else the wrapper is not installed and the engine behaves as before. Validation: - Configures and builds in full on this host: Ninja, MSVC 19.44.35207, CUDA 12.4.131, `-DCMAKE_CUDA_ARCHITECTURES=60-real;75-real -DSTRATA_EXPERIMENTAL_SM60=ON -DSTRATA_BUILD_TESTS=ON`, linking `strata.exe` straight from `fb58e0d`. - Based on 0.1.41 (`fb58e0d`), three files, +222/-1: two new files plus one block in `src/program/generate.cpp`. The added lines and both new files are byte-identical to the tree the full-model runs below were measured on. - Windows, CUDA 12.4.131, SM75 + SM60, RTX 2080 Ti + Tesla P100, driver 537.13: the integrated engine build passes, and every request of the fixtures below returned its expected JSON with `cache_n = 0`. - Same-day A/B on an 11-request fixture (422 - 5127 token prompts, four of them cold): the engine's own prompt time went from 91.5 s to 65.7 s (-28.3%), and 92.1 GB (64.8%) fewer bytes were read from the files (142.1 GB -> 50.1 GB, disk counters per request). - The source's own contribution, with `--short-read` held equal: an 884-token prompt reads in 7.12 s from the files (8.06 ms/token) and in 4.87 s with a helper's cache (5.50 ms/token). Crossing from the verify windows to the batched path at the measured crossover adds a further 1.15 s on that request. - With `--short-read 768` the same suite is 59.0 s; the two 884-token requests go from 8.41 s to 4.91 s each. That flag is a configuration choice and is not part of this patch. Limits: only the serial serve path is covered; the layer split, batch slots, HIP and non-Windows builds keep the old source. Nothing changes when a helper does not hold the blob, when no helper GPU is present, or when the helper holds a different placement of the experts. The measurements are one Windows dual-GPU host with the native IQ3_XXS pack at 8192 context; the flag is not proposed as a default.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.