Pull requests / #853
#853 cuda: specialize hyper-connection up for short verify windows
closed · @InB4DevOps · 0 Kommentare · Auf GitHub
BenchmarksNVIDIA / CUDAModels & quantsWindows
Beschreibung
## Summary Specialize the multi-token hyper-connection up kernel for verify windows with four or fewer tokens. The existing kernel always reserved shared memory and unrolled work for the maximum eight tokens. This change templates the token capacity and dispatches a four-token specialization for `T <= 4`. Larger windows continue to use the existing eight-token kernel. `STRATA_GR_UP_MAX4=0` restores the previous behavior for regression testing. ## Correctness The specialization preserves launch geometry, arithmetic, and accumulation order. `gr_parity` now checks direct execution and graph replay at `T=2`, `T=4`, and `T=8`. The optimized kernel was bit-identical to the existing path, and end-to-end greedy responses were identical in all A/B runs. ## Measurements Measured on an RTX 3060 12 GB with the Coder IQ1_M native pack, `--spec 4`, and `--spec-min-p 0.70`. First A/B: | | Baseline | Specialized | |---|---:|---:| | Decode | 42.08 ms/window | 41.33 ms/window | | GPU total | 36.43 ms/window | 35.73 ms/window | | Throughput | 38.3 tok/s | 39.0 tok/s | Repeat: | | Baseline | Specialized | |---|---:|---:| | Decode | 42.26 ms/window | 41.38 ms/window | | GPU total | 36.61 ms/window | 35.76 ms/window | | Throughput | 38.1 tok/s | 38.9 tok/s | After enabling the specialization by default, a final default-versus-legacy run measured: - Default: 41.33 ms/window, 39.0 tok/s - Legacy (`STRATA_GR_UP_MAX4=0`): 41.89 ms/window, 38.4 tok/s - Identical tokens and speculative acceptance The savings appeared in both hyper-connection reads, while unrelated stages remained effectively unchanged. ## Validation - Built `strata` and `gr_parity` - `gr_parity --selftest` - End-to-end greedy A/B runs with identical tokens and acceptance - Default-versus-legacy regression run
Mehr auf der Site
Links zu Install, Modellen, Releases.