Pull requests / #1724

#1724 engine: MXFP4 routed experts (MXFP4_MOE GGUFs) and MXFP4 PLE tables

open · @VLOD-ZDOV · 0 コメント · GitHub で見る

Setup & installAMD / HIPNVIDIA / CUDAModels & quantsWindows

本文

## Title
engine: MXFP4 routed experts (MXFP4_MOE GGUFs) and MXFP4 PLE tables

## Summary
GGUFs made with llama-quantize's `MXFP4_MOE` keep every routed expert (gate, up and down) and the PLE table in MXFP4
(ggml type 39). The engine refused them at start: `layer 0's experts are MXFP4/MXFP4 (ggml types 39/39), which this
engine has no GPU kernels for`. With this change such a file runs as a native pack (`tools/iq_pack.py --compat-bf16`,
unchanged). The formats that already run take the same kernels as before (checked below).

## What changed
- `src/kernels/cuda/iq_kernels.cu`
  - MXFP4 as an expert format for gate/up and down: `Fmt<39>` and `Split<39>`. It is IQ4_NL's layout with the FP4
    table (`kvalues_mxfp4`), one E8M0 exponent per 32 values and byte-aligned nibbles (17-byte blocks), after
    llama.cpp's `vec_dot_mxfp4_q8_1`.
  - The prompt path's dequantizer (`dq_mxfp4`), the row size and `is_iq`.
  - **A fix the new format exposed:** `native_gu_multi_kernel` took its 16-lane path (`row_dot_80_sub16`, which
    covers 80 dot calls of a row) whenever `n_embd == 2560`. That held for every split gate/up format so far
    (256-value blocks: 10 x 8 = 80), but an MXFP4 row is 80 x 2 = 160 calls, so only half of each gate/up row was
    computed. The kernel and the host's grid now test the K itself (`nb * ipb == 80`). For the existing formats the
    condition is the same.
- `include/strata/artifact/gguf_reader.hpp`: MXFP4's block geometry (32 values in 17 bytes).
- `include/strata/artifact/dequant.hpp`: `dequantize_mxfp4`, after ggml's `dequantize_row_mxfp4` and
  `GGML_E8M0_TO_FP32_HALF`.
- `include/strata/kernels/ngram.hpp`, `src/kernels/ngram.cpp`: MXFP4 as a PLE table format (85-byte rows).
- Tests:
  - `src/kernels/iq_parity.cpp`, `tools/iq_fixture.py`: an MXFP4 fixture (E8M0 scales, including the two subnormal
    codes); a type without a dense native MMVQ goes through `iq_mmvq`, the experts' own `Fmt<T>::dot`.
  - `src/kernels/native_grouped_parity.cpp`: MXFP4 in the format lists, and `check_old`, which compares the default
    kernels with the one-row kernels (`iq_set_old_kernels`) at the model's shapes (2560 x 640). The v1 comparison
    could not catch the bug above, because both sides take the same gate/up kernel.
  - `src/ngram/ple_reader_test.cpp`: MXFP4 in the test's type table.

Not changed: the prompt path's MMQ. ggml's MXFP4 MMQ on Blackwell (sm_120a) uses FP4 tensor-core instructions that
expect FP4-quantized activations, and Strata hands MMQ q8_1 activations, so MXFP4 layers take the existing FP16
fallback (dequantize + GEMM). setup.py still names MXFP4 files as unsupported; that is left for a separate change.

## Extra Notes
How it was checked (CUDA build only):
- `iq_parity`: MXFP4 dequant rel 0.00e+00 against gguf-py; MMVQ rel 5.48e-03 (the other formats: 4.6e-03 to
  5.7e-03); multi-column calls bitwise equal to one-column calls.
- `native_grouped_parity`: 0 failures. MXFP4 with every format pair is bitwise equal to v1. `check_old`: IQ3_XXS/IQ4_NL,
  IQ3_S/IQ4_NL and MXFP4/MXFP4 at 2560 x 640 are bitwise equal to the one-row kernels (rel 0.00e+00). With the old
  `xb == 80` condition put back, `check_old` fails MXFP4 (rel 8.7e-01), so the check covers the bug.
- `ple_reader_test --selftest`: every format including MXFP4 OK. `dequantize_mxfp4` is bitwise equal to gguf-py on
  six real rows of an MXFP4 PLE table, the first and last rows included.
- A real MXFP4_MOE GGUF of Qwen3.8-Flash-Next answers correctly and coherently. Before the fix it looped on the same
  paragraph.

Not tested: the HIP and SYCL builds (`iq_kernels.cu` is also compiled for HIP), Windows, and other GPU generations.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

関連リンク

インストール・モデル・リリースへの站内リンク。