Pull requests / #764
#764 decode: STRATA_ADAPT_LAG - #463's reproducibility without its stall
closed · @merbanan · 0 comments · View on GitHub
Server & APINVIDIA / CUDAModels & quantsWindows
Description
#463 made a verify window wait for the adaptive tier's previous swaps (cudaEventSynchronize) so that whether a swapped-in expert runs on the GPU or the CPU in a window - they round differently - no longer depends on when its PCIe copy lands. The wait is bounded by the copy volume: on an RTX 2060 SUPER 8 GB over PCIe 3.0 x8 (Q2_0, 2,262 cache slots) it costs ~4.6 ms of a ~50 ms window at every context length, the same as dropping it with STRATA_ADAPT_NOWAIT=1 recovers. Now a window takes the pending swaps only once they are STRATA_ADAPT_LAG windows old (default 2), and then waits for their copies, which have landed by then. The window that first computes a swapped-in expert on the GPU still depends only on the window count, so greedy output is as reproducible as with #463's wait; a swapped-in expert just counts one window later. STRATA_ADAPT_LAG=1 is #463's behaviour; STRATA_ADAPT_NOWAIT=1 keeps its meaning. Both decode loops (serve and generate).
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.