Pull requests / #1720
#1720 dequant: coalesced Q8_0 path (four threads per block, 16-byte stores), bit-identical, 7x faster on V100
open · @eelgaev · 0 commentaires · Sur GitHub
BenchmarksNVIDIA / CUDAModels & quants
Description
The prompt path dequantizes each dense Q8_0 projection (GDN qkv/gate/out, QSA q/k/v/out, shared expert, hc) into the GEMM scratch before cuBLAS. `dequant_kernel` runs one thread per 32-value block, so a warp's 2-byte stores land 64 bytes apart: about 70 GB/s of writes on a V100. `dequant_q8_0_kernel` gives each block to four threads, 8 values each, written with 16-byte stores: a warp reads 8 consecutive blocks and writes 512 consecutive bytes. The arithmetic per value is unchanged (`d * q`), so the output is the same bits. It is used for Q8_0 only when the output rows are contiguous (`ld == cols`) and the pointers are aligned (blocks to 2 bytes, output to 16); everything else (including the padded `ld` scratch) takes `dequant_kernel` as before. One file, 32 lines added, nothing else changed. ### Bitwise Every 2-D Q8_0 tensor of Qwen3.8-Flash-Next UD-Q4_K_XL (498 tensors) through `dequant_f32`, `dequant_bf16` and `dequant_f16`, built from `main` (fb58e0d) and from this branch: an FNV-1a hash over all outputs is equal for all three output types. `dequant_bf16_test` covers Q8_0 against the CPU reference as before. ### Speed (V100-SXM2-16GB, each call timed on its own) | | before | after | |---|---:|---:| | attn_qkv 10240 x 2560, FP16 | 0.740 ms (70 GB/s written) | **0.106 ms** (495 GB/s) | | attn_q 12288 x 2560, FP16 | 0.887 ms | **0.126 ms** | | ssm_out 2560 x 6144, FP16 | 0.451 ms | **0.065 ms** | | all 498 tensors, FP16 | 138.5 ms | **20.4 ms** | | all 498 tensors, BF16 | 140.8 ms | **20.5 ms** | | all 498 tensors, FP32 | 190.9 ms | **35.4 ms** | Whole prompts on 4x V100, 4,096-token chunks (where this was first used): 18K 4,537 -> 4,708, 65K 6,540 -> 6,766, 123K 6,783 -> 6,985 tok/s (+3-4%). The kernel is plain CUDA (no architecture-specific instructions), so newer cards take the same path.
Sur le site
Liens install, modèles, releases.