Pull requests / #1653

#1653 [WIP] cuda: SM75 IQ4_XS four-row reuse and NW3 shape specialization

open · draft · @Unmaple · 0 comentários · No GitHub

BenchmarksAMD / HIPNVIDIA / CUDADocumentation

Descrição

WIP / 施工中. Not ready for merge or review.

For IQ4_XS native/exact projections, reuse each Q8_1 activation fragment across four output rows and optionally remove the empty fourth warp for the measured K=2560 shapes. This is opt-in, SM75-only, T=2-4; default, TSUM/non-exact, HIP and unsupported calls retain their old dispatch.

Related: #1418 already implements general row reuse. This draft focuses on the additional NW3 shape specialization and its qualification; the shared mechanism must be consolidated before merging.

Historical affected projection improvements were 7.622 +/- 0.484% (QSA) and 5.585 +/- 0.660% (GDN), incremental over four-row NW4. These are GPU-stage measurements, not tok/s gains. Full scope, CI method, same-input checks and pending tests: [report](docs/SM75_IQ4_REUSE_WIP.md).

CUDA SM75 translation-unit compilation passed. The expanded synthetic parity harness passed: 980,113 exact-layout outputs, zero bit differences/non-finite outputs; a powered negative control found differences. Full integration, current-head real-input parity, HIP build and production end-to-end measurements remain pending. Defaults are not changed.

No site

Links install, modelos, releases.