Pull requests / #478
#478 prefill: MMQ prompt kernels for Q6_K experts (STRATA_MMQ_KQUANTS)
closed · draft · @yannickloth · 0 コメント · GitHub で見る
AMD / HIPNVIDIA / CUDAModels & quants
本文
`STRATA_MMQ_KQUANTS` builds the CUDA MMQ prompt kernels for Q4_K, Q5_K and Q5_1. Q6_K is the same K-quant family and ggml already ships an MMQ instance for it (`ggml/src/ggml-cuda/template-instances/mmq-instance-q6_k.cu`), so experts stored as Q6_K currently fall back to dequant-to-FP16 + cuBLAS on the prompt path. This wires Q6_K through the same build option: - `CMakeLists.txt`: append `q6_k` to `_strata_mmq_cuda_types` under `STRATA_MMQ_KQUANTS`; option comment/description now say Q4_K / Q5_K / Q5_1 / Q6_K. - `src/prefill/moe_mmq.cu`: `GGML_TYPE_Q6_K` in `supported()` and the `mul_mat_q_case` dispatch, both inside the existing `#ifdef STRATA_MMQ_KQUANTS`. - `tests/cuda/prefill_mmq_kquant_test.cpp`: screen Q6_K gate/up (1280 x 2560) against the dequantized reference, and drop the `"Q6_K should not be covered"` assertion (it is covered now). The option stays off by default, for the same reason as the others: the kernels load eagerly and cost VRAM / expert slots. **Status — please read before merging (opened as a draft):** - I could not build CUDA where this was prepared. The change is mechanical and the Q6_K MMQ instance / `DECL_MMQ_CASE(GGML_TYPE_Q6_K)` exist in the pinned llama.cpp, but this needs a build of `prefill_mmq_kquant_test` (`STRATA_MMQ_KQUANTS=ON`, `STRATA_BUILD_TESTS=ON`) and CI. - No A3000 prompt-speed numbers for a Q6_K pack are attached yet; the sm_86 figures in the fork are for the Q4_K/Q5_K/Q5_1 option. - Design question: I extended the existing option rather than adding a new one. If you'd rather keep `STRATA_MMQ_KQUANTS` scoped to UD-Q4_K_XL's formats, Q6_K can move behind its own option instead.
関連リンク
インストール・モデル・リリースへの站内リンク。