Pull requests / #203
#203 serve: prompt checkpoints copied out asynchronously (pinned staging on the prompt stream + a thread)
closed · @q8atnight · 0 comments · View on GitHub
Server & APIMulti-GPUNVIDIA / CUDAModels & quants
Description
## Summary A mid-prompt checkpoint (every `--prompt-cache-every` tokens, default 16K) currently stops the GPU: `cudaDeviceSynchronize()` plus pageable `cudaMemcpy`s into fresh vectors - ~370 ms of GPU idle per checkpoint on our box. This copies the same bytes out asynchronously: the running state goes to a pinned staging buffer through stream-ordered `cudaMemcpyAsync`s on the prompt stream (ordered after the chunk that produced the state, before the next chunk overwrites it; ~10 ms of copy engine), and a thread moves it into the checkpoint's own vectors while the next chunk runs. The checkpoint joins `checks` at the next checkpoint, at a turn/root checkpoint, or once the prompt is read. One commit on top of 0.1.27 (a790805). ## Measured 2x RTX 3090, IQ3_S, developed in the 0.1.24 + dual-GPU engine: the synchronous copy-out held the GPU idle ~370 ms per mid-prompt checkpoint; async it overlaps the next chunk, so the prompt path does not wait for it. With `--prompt-cache-every 16384` a 128K prompt takes ~7 checkpoints, so up to ~2.6 s of GPU idle (7 x ~370 ms, arithmetic, not separately measured). ## Correctness The checkpoint contents are unchanged - the same bytes, only copied later. The copy is stream-ordered between the chunk that wrote the running state and the chunk that overwrites it, so no sync is needed on the prompt path. A layer split keeps upstream's path untouched: its checkpoints are assembled from the stages' own parts, saved on their own streams, and a per-stage pinned staging does not obviously fit there - say so if you want it extended. ## Switch `STRATA_CK_SYNC=1` = the old synchronous path. `STRATA_PREFILL_TIMING=1` traces the staging allocation and the enqueue time of each async checkpoint.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.