Pull requests / #284

#284 Verify: the commit graph no longer waits; it overlaps the MTP draft

closed · @sergqwer · 0 comentarios · En GitHub

BenchmarksServer & APINVIDIA / CUDAWindows

Descripción

## Summary

The verify window's commit ended in `cudaStreamSynchronize`, which held the host ~0.45 ms a round. The next window
follows the commit on the same stream, and the drafter (on its own stream) reads only this window's final rows and
its own K/V. So on a single GPU the commit now returns at once (`Verifier::set_commit_async`).

Whatever touches the session from another stream or from the host afterwards synchronizes the device first: the
start of each `--serve` request (checkpoint restore, zeroing, the prompt path) and the end of a generate run.
Checkpoint saves already did. A layer split keeps the wait.

## Measured

Measured on current main (0.1.29, d6708a4) with ISTA-DASLab's IQ2_XS, RTX 5090 32 GB, Ryzen 9 9950X3D, 128 GB DDR5, Windows 11. KL is first-token KL divergence (`STRATA_DUMP_FIRST_LOGITS`, #276); two runs of main itself differ by 0 on 5 of 6 short prompts and by 0.00013 on the sixth. 400-token answers, `--vram-reserve-mib 1500` (#279), 3 runs each,
interleaved:

| | decode |
| --- | ---: |
| main | 139.8 / 123.0 / 122.9 tok/s (mean 128.6) |
| this PR | 139.0 / 130.6 / 137.8 tok/s (mean 135.8) |

That is +5.6%. The generated tokens are identical in every pair. On an NVFP4 pack, which has more GPU work per
round, the gain was +0.8% (114.7 -> 115.6 tok/s), within the noise. Five consecutive `--serve` requests over two
conversations (fork) answered correctly, with the same prompt reuse as with the wait.

I know #203 was closed for showing no difference. This one is small too, so I'm leaving the call to you.

## Switch

`STRATA_COMMIT_SYNC=1` keeps the wait.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

En el sitio

Enlaces a install, modelos, releases.