Pull requests / #599
#599 native experts: Q4_0 and Q4_1 GPU kernels, Q4_0 PLE table
closed · @saikiran-rs · 0 comments · View on GitHub
BenchmarksNVIDIA / CUDAModels & quantsDocumentation
Description
A plain llama-quantize Q4_0 file of this model (bartowski's Qwen3.8-Flash-Next-Q4_0 and Swift-1.5 Q4_0: gate/up Q4_0, down Q4_0 on 42 layers and Q4_1 on 6, a Q4_0 PLE table and token embedding) was refused at start: no GPU kernels for Q4_0/Q4_1 experts, and the PLE reader took only IQ4_NL, Q5_0 and FP8. Follows #222 (Q4 to Q6 quantizations). - Q4_0: Fmt<2> with the codes centred before dp4a (q - 8 as signed bytes), the exact integer form ggml-cpu's q4_0 x q8_0 dot uses, rather than llama.cpp's CUDA chain (uncentred codes minus 8 x the q8_1 float sum): on a real Q4_0 expert the CUDA chain measured 1.4e-2 relative error against the float product, the centred form 5.3e-3. - Q4_1: Fmt<3>, the Q5_1 chain without the fifth bit and the same min-term choice. - dq_q4_0 / dq_q4_1 for the prompt path's dequant fallback; row bytes, is_iq and the GU/D/MMVQ format lists. - PLE: Q4_0 rows (90 B, as IQ4_NL) through the same reader. - docs/UNSLOTH_Q4.md: a Q4_0 section with the measured run. native_expert_parity --synthetic (now in the CMake matrix): q4_0/q4_0 CPU 1.29e-2, GPU 1.29e-2, CPU-GPU 1.1e-7; q4_0/q4_1 CPU-GPU 3.2e-4; GPU dequant bit-exact against to_float. The existing pairs are unchanged. On bartowski's file (revision 928589f), layers 0 and 5 (Q4_1 down) and 20 and 47: CPU and GPU 1.1-1.3e-2, CPU-GPU 1.1e-7 (Q4_0 down) and 5e-4 to 1.1e-3 (Q4_1 down). Served on one RTX 3060 (all experts in RAM, 131K context): 24,461-token prompt 787 tok/s read, 24.0 tok/s out; 67,934 tokens 737 / 26.4, the hidden fact found.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.