Pull requests / #820
#820 prefill: the dense GGUF projections through MMQ on HIP (opt-in STRATA_DENSE_MMQ=1)
closed · @xendak · 0 commentaires · Sur GitHub
BenchmarksAMD / HIPNVIDIA / CUDAModels & quants
Description
## Notes
This repository is being on the move constantly so its hard to catch up to the
`current` branch, but i'll be trying to keep this up to date.. i don't have an
ai subscription though so sometimes the update speed is too demanding :')
## What
The prompt path's dense GGUF projections run dequantize-to-FP16 + cuBLAS. On the
AMD cards (checked with rocprof on an RX 6900 XT) that is 68% of the prefill GPU
time at ~25% of the card's fp32-accumulate rate, while the same shapes through
llama.cpp's int8 MMQ run ~7x faster. This PR wires them through the MMQ path the
MoE experts already take, opt-in:
- `Gemm::native`'s `beta == 0` products with a covered type, a K that is a
multiple of 256 values, and a matrix whose tile fits the card's shared memory
run `mmq::quantize` + one `mul_mat_q` per 1024-row chunk (the FP16 activations
are widened to FP32 in the dequant scratch, which the old path only fills
afterwards). **`STRATA_DENSE_MMQ=1` opts in; unset, nothing changes.** The
first dense call prints `prefill gemm: dense MMQ on`.
- `STRATA_MMQ_KQUANTS` now also compiles the `q4_k q5_k q5_1 q6_k` instances and
defines `STRATA_MMQ_KQUANTS` on the HIP branch (the mixed-quant packs' dense
tensors are K-quants: this model's attention/PLE matrices are
IQ4_XS/IQ3_S/Q4_K/Q5_K/Q6_K), and adds `q6_k` to the CUDA list (the usual
attention weight of plain i-quant GGUFs).
- `moe_mmq` grows two small helpers: `f16_to_f32` and a `{0, rows}` bounds
writer.
- `hip_prefill_mmq_parity` gains synthetic Q4_K/Q5_K/Q6_K/Q5_1 passes (skipped
when the build does not have the instances).
## Why K must be a full 256-value multiple
llama.cpp's MMQ loads the weights in 256-value K chunks; the chunk past a
partial row reads bytes past the row - the next row's, or, for the last weight
row, **past the whole tensor**. The MoE path is safe because its expert gather
buffers end in a zeroed tail; a dense GGUF tensor is passed as-is, so what
follows it is the next tensor's blocks in the mapping. On swift-1.5-iq3_xxs the
640-value IQ4_NL shared-expert down picked a NaN scale out of those bytes and
produced a whole NaN output column, which zeroed the model output - found and
minimized with a type-bisect env knob and an operand-capture replay harness. The
odd shapes stay on the dequantize + cuBLAS path.
## Testing (RX 6900 XT, gfx1030, ROCm 7.2.3, i7-13700KF)
`-DSTRATA_MMQ_KQUANTS=ON -DSTRATA_ENABLE_HIP=ON`, swift-1.5-iq3_xxs, the user's
knobs
(`--pool-workers 15 --adapt-every 1 --prefill 8192 --kv-resident 32768 --vram-reserve-mib 1024`),
`bench_prefill.py` fresh medians of two ~4K and two ~8K trials:
| | 4K fresh | 8K fresh | 8K followup | decode |
| ---------------------- | --------------- | --------------- | --------------- | ---------- |
| STRATA_DENSE_MMQ unset | 436.1 tok/s | 447.7 tok/s | 163.9 tok/s | 40.6 tok/s |
| STRATA_DENSE_MMQ=1 | **757.2 tok/s** | **773.9 tok/s** | **206.1 tok/s** | 40.3 tok/s |
+73-74% fresh prefill; decode within noise. Greedy output is bit-identical for
the first 200 tokens of a short prompt; longer greedy runs diverge around token
~26 (the q8_1 activation rounding) and stay coherent. `hip_prefill_mmq_parity`
passes with the new K-quant passes (rel_l2 ≤ 0.005 vs ggml's own dequantizer on
the host).
Not tested: CUDA builds (the instance list change is mechanical), other AMD
archs, the q6_k CUDA instance end to end.
Sur le site
Liens install, modèles, releases.