Pull requests / #478

#478 prefill: MMQ prompt kernels for Q6_K experts (STRATA_MMQ_KQUANTS)

closed · draft · @yannickloth · 0 Kommentare · Auf GitHub

AMD / HIPNVIDIA / CUDAModels & quants

Beschreibung

`STRATA_MMQ_KQUANTS` builds the CUDA MMQ prompt kernels for Q4_K, Q5_K and Q5_1. Q6_K is the same K-quant family and ggml already ships an MMQ instance for it (`ggml/src/ggml-cuda/template-instances/mmq-instance-q6_k.cu`), so experts stored as Q6_K currently fall back to dequant-to-FP16 + cuBLAS on the prompt path.

This wires Q6_K through the same build option:

- `CMakeLists.txt`: append `q6_k` to `_strata_mmq_cuda_types` under `STRATA_MMQ_KQUANTS`; option comment/description now say Q4_K / Q5_K / Q5_1 / Q6_K.
- `src/prefill/moe_mmq.cu`: `GGML_TYPE_Q6_K` in `supported()` and the `mul_mat_q_case` dispatch, both inside the existing `#ifdef STRATA_MMQ_KQUANTS`.
- `tests/cuda/prefill_mmq_kquant_test.cpp`: screen Q6_K gate/up (1280 x 2560) against the dequantized reference, and drop the `"Q6_K should not be covered"` assertion (it is covered now).

The option stays off by default, for the same reason as the others: the kernels load eagerly and cost VRAM / expert slots.

**Status — please read before merging (opened as a draft):**

- I could not build CUDA where this was prepared. The change is mechanical and the Q6_K MMQ instance / `DECL_MMQ_CASE(GGML_TYPE_Q6_K)` exist in the pinned llama.cpp, but this needs a build of `prefill_mmq_kquant_test` (`STRATA_MMQ_KQUANTS=ON`, `STRATA_BUILD_TESTS=ON`) and CI.
- No A3000 prompt-speed numbers for a Q6_K pack are attached yet; the sm_86 figures in the fork are for the Q4_K/Q5_K/Q5_1 option.
- Design question: I extended the existing option rather than adding a new one. If you'd rather keep `STRATA_MMQ_KQUANTS` scoped to UD-Q4_K_XL's formats, Q6_K can move behind its own option instead.

Mehr auf der Site

Links zu Install, Modellen, Releases.