Pull requests / #284
#284 Verify: the commit graph no longer waits; it overlaps the MTP draft
closed · @sergqwer · 0 comentarios · En GitHub
BenchmarksServer & APINVIDIA / CUDAWindows
Descripción
## Summary The verify window's commit ended in `cudaStreamSynchronize`, which held the host ~0.45 ms a round. The next window follows the commit on the same stream, and the drafter (on its own stream) reads only this window's final rows and its own K/V. So on a single GPU the commit now returns at once (`Verifier::set_commit_async`). Whatever touches the session from another stream or from the host afterwards synchronizes the device first: the start of each `--serve` request (checkpoint restore, zeroing, the prompt path) and the end of a generate run. Checkpoint saves already did. A layer split keeps the wait. ## Measured Measured on current main (0.1.29, d6708a4) with ISTA-DASLab's IQ2_XS, RTX 5090 32 GB, Ryzen 9 9950X3D, 128 GB DDR5, Windows 11. KL is first-token KL divergence (`STRATA_DUMP_FIRST_LOGITS`, #276); two runs of main itself differ by 0 on 5 of 6 short prompts and by 0.00013 on the sixth. 400-token answers, `--vram-reserve-mib 1500` (#279), 3 runs each, interleaved: | | decode | | --- | ---: | | main | 139.8 / 123.0 / 122.9 tok/s (mean 128.6) | | this PR | 139.0 / 130.6 / 137.8 tok/s (mean 135.8) | That is +5.6%. The generated tokens are identical in every pair. On an NVFP4 pack, which has more GPU work per round, the gain was +0.8% (114.7 -> 115.6 tok/s), within the noise. Five consecutive `--serve` requests over two conversations (fork) answered correctly, with the same prompt reuse as with the wait. I know #203 was closed for showing no difference. This one is small too, so I'm leaving the call to you. ## Switch `STRATA_COMMIT_SYNC=1` keeps the wait. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
En el sitio
Enlaces a install, modelos, releases.