Pull requests / #203

#203 serve: prompt checkpoints copied out asynchronously (pinned staging on the prompt stream + a thread)

closed · @q8atnight · 0 comentarios · En GitHub

Server & APIMulti-GPUNVIDIA / CUDAModels & quants

Descripción

## Summary

A mid-prompt checkpoint (every `--prompt-cache-every` tokens, default 16K) currently stops the GPU:
`cudaDeviceSynchronize()` plus pageable `cudaMemcpy`s into fresh vectors - ~370 ms of GPU idle per checkpoint on
our box. This copies the same bytes out asynchronously: the running state goes to a pinned staging buffer through
stream-ordered `cudaMemcpyAsync`s on the prompt stream (ordered after the chunk that produced the state, before
the next chunk overwrites it; ~10 ms of copy engine), and a thread moves it into the checkpoint's own vectors
while the next chunk runs. The checkpoint joins `checks` at the next checkpoint, at a turn/root checkpoint, or
once the prompt is read. One commit on top of 0.1.27 (a790805).

## Measured

2x RTX 3090, IQ3_S, developed in the 0.1.24 + dual-GPU engine: the synchronous copy-out held the GPU idle
~370 ms per mid-prompt checkpoint; async it overlaps the next chunk, so the prompt path does not wait for it.
With `--prompt-cache-every 16384` a 128K prompt takes ~7 checkpoints, so up to ~2.6 s of GPU idle (7 x ~370 ms, arithmetic, not separately measured).

## Correctness

The checkpoint contents are unchanged - the same bytes, only copied later. The copy is stream-ordered between the
chunk that wrote the running state and the chunk that overwrites it, so no sync is needed on the prompt path.
A layer split keeps upstream's path untouched: its checkpoints are assembled from the stages' own parts, saved on
their own streams, and a per-stage pinned staging does not obviously fit there - say so if you want it extended.

## Switch

`STRATA_CK_SYNC=1` = the old synchronous path. `STRATA_PREFILL_TIMING=1` traces the staging allocation and the
enqueue time of each async checkpoint.

En el sitio

Enlaces a install, modelos, releases.