Pull requests / #1653
#1653 [WIP] cuda: SM75 IQ4_XS four-row reuse and NW3 shape specialization
open · draft · @Unmaple · 0 commentaires · Sur GitHub
BenchmarksAMD / HIPNVIDIA / CUDADocumentation
Description
WIP / 施工中. Not ready for merge or review. For IQ4_XS native/exact projections, reuse each Q8_1 activation fragment across four output rows and optionally remove the empty fourth warp for the measured K=2560 shapes. This is opt-in, SM75-only, T=2-4; default, TSUM/non-exact, HIP and unsupported calls retain their old dispatch. Related: #1418 already implements general row reuse. This draft focuses on the additional NW3 shape specialization and its qualification; the shared mechanism must be consolidated before merging. Historical affected projection improvements were 7.622 +/- 0.484% (QSA) and 5.585 +/- 0.660% (GDN), incremental over four-row NW4. These are GPU-stage measurements, not tok/s gains. Full scope, CI method, same-input checks and pending tests: [report](docs/SM75_IQ4_REUSE_WIP.md). CUDA SM75 translation-unit compilation passed. The expanded synthetic parity harness passed: 980,113 exact-layout outputs, zero bit differences/non-finite outputs; a powered negative control found differences. Full integration, current-head real-input parity, HIP build and production end-to-end measurements remain pending. Defaults are not changed.
Sur le site
Liens install, modèles, releases.