Pull requests / #853

#853 cuda: specialize hyper-connection up for short verify windows

closed · @InB4DevOps · 0 评论 · 在 GitHub 查看

BenchmarksNVIDIA / CUDAModels & quantsWindows

描述

## Summary

Specialize the multi-token hyper-connection up kernel for verify windows with four or fewer tokens.

The existing kernel always reserved shared memory and unrolled work for the maximum eight tokens. This change templates the token capacity and dispatches a four-token specialization for `T <= 4`. Larger windows continue to use the existing eight-token kernel.

`STRATA_GR_UP_MAX4=0` restores the previous behavior for regression testing.

## Correctness

The specialization preserves launch geometry, arithmetic, and accumulation order.

`gr_parity` now checks direct execution and graph replay at `T=2`, `T=4`, and `T=8`. The optimized kernel was bit-identical to the existing path, and end-to-end greedy responses were identical in all A/B runs.

## Measurements

Measured on an RTX 3060 12 GB with the Coder IQ1_M native pack, `--spec 4`, and `--spec-min-p 0.70`.

First A/B:

| | Baseline | Specialized |
|---|---:|---:|
| Decode | 42.08 ms/window | 41.33 ms/window |
| GPU total | 36.43 ms/window | 35.73 ms/window |
| Throughput | 38.3 tok/s | 39.0 tok/s |

Repeat:

| | Baseline | Specialized |
|---|---:|---:|
| Decode | 42.26 ms/window | 41.38 ms/window |
| GPU total | 36.61 ms/window | 35.76 ms/window |
| Throughput | 38.1 tok/s | 38.9 tok/s |

After enabling the specialization by default, a final default-versus-legacy run measured:

- Default: 41.33 ms/window, 39.0 tok/s
- Legacy (`STRATA_GR_UP_MAX4=0`): 41.89 ms/window, 38.4 tok/s
- Identical tokens and speculative acceptance

The savings appeared in both hyper-connection reads, while unrelated stages remained effectively unchanged.

## Validation

- Built `strata` and `gr_parity`
- `gr_parity --selftest`
- End-to-end greedy A/B runs with identical tokens and acceptance
- Default-versus-legacy regression run

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。