Pull requests / #292

#292 NVFP4 routed experts: converter, pack, decode, prompt path (W4A8, W4A4 on Blackwell), AVX-512 CPU rows

closed · @sergqwer · 0 comments · View on GitHub

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows

Description

## Summary

A ModelOpt NVFP4 checkpoint of Qwen3.8-Flash-Next runs on the native path: routed experts in NVFP4 (group 16), the
rest in the formats the native path already reads. `docs/NVFP4.md` has the whole design and every measurement.
These changes come from an NVFP4 fork of Strata (sergqwer/strata-nvfp4), ported onto current main one topic per PR.

**Pipeline.**
- **Converter** (`tools/nvfp4_convert.py`): llama.cpp's own converter (the pinned commit) with the engine's type
  policy. The experts are repacked without loss; `tools/nvfp4_verify.py` and `nvfp4_verify_gguf.py` compare them with
  the checkpoint bit for bit.
- **Pack** (`tools/iq_pack.py`): each NVFP4 blob gets a 16-byte tail `{s_gate, s_up, s_down, 0}`, the per-expert
  `weight_scale_2`. Every consumer that already moves whole blobs (arena, device cache, prefill staging, a second
  GPU) then moves the scales with the weights. `experts.bin` is written by itself, because it is the only source with
  the tails.
- **Decode**: NVFP4 in the native expert kernels (q8_1 activations, UE4M3/E2M1 decoded in software, bit-identical to
  ggml). Each global scale multiplies its own projection's FP32 output: folding `s_down` into up put the hidden into
  FP16's subnormals (2-12% expert error instead of 1.1%).
- **CPU pool**: AVX-512 NVFP4 rows, 1.8-3.7x faster than ggml-cpu's (identical arithmetic, another summation order);
  ggml-cpu without AVX-512.
- **Prompt path**: llama.cpp MMQ. The default is **W4A8**: the int8 MMQ, compiled for sm_120 with Blackwell's FP4 MMA
  hidden (`mmq_nvfp4_w4a8.cu`). **W4A4** (FP4 MMA, sm_120a) and **FP16** are opt-in through `STRATA_PREFILL_NVFP4`.
- **PCIe share**: 0.25 of the missed experts for NVFP4 packs, not 0.55. Their blobs are 2.76 MB, and at 0.55 the copy
  kernel was 40% of GPU time.
- **Build**: CMake turns 120 into 120a (as ggml's own CMake does), adds the NVFP4 MMQ instance and the parity tools,
  and guards the CUDA-only parts for HIP.

## Measured

RTX 5090, Ryzen 9 9950X3D, 128 GB, Windows 11; this branch on main, 64K context, int8 K/V, `--vram-reserve-mib 1500`:

| | NVFP4 pack | IQ2_XS, same engine |
| --- | ---: | ---: |
| experts | 63.3 GiB | 33.0 GiB |
| expert cache slots | ~8,350 | ~17,300 |
| decode, 400-token answers | 114-116 tok/s | 132-139 tok/s |
| a 32K prompt | ~3,900 tok/s | ~5,600 tok/s |

Precision of the prompt path, first-token KL against FP16, 8 prompts:

| mode | KL mean | top-1 |
| --- | ---: | ---: |
| W4A8 (default) | 0.0018 | 8/8 |
| W4A4 | 0.0080 | 7/8 |
| noise floor | 0.00023 | 8/8 |

W4A8 reads prompts at 2,503 tok/s, W4A4 at 2,904 and FP16 at 1,905 (4-8K prompts, 262K context).

## Tests

- `nvfp4_expert_gpu_parity`: the decode kernels against ggml's reference.
- `nvfp4_avx512_parity`: the CPU rows against ggml-cpu.
- `mmq_nvfp4_parity`: the prompt path's products against FP64.
- The verify scripts: the conversion, bit for bit.
- sm_75/86/89 builds (emulated on the RTX 5090): the parity tools give results identical to sm_120a.

## Tested with, and what is not in this PR

- **Checkpoints.** Run end to end with `jpezzulli/OrcaRouter-Qwen3.8-Flash-Next-Uncensored-ModelOpt-NVFP4`.
  `nvidia/Qwen3.8-Flash-Next-NVFP4` quantizes its routed experts the same way (by its `hf_quant_config.json`), so it
  takes the same path. I have not run it.
- **Not included:**
  - setup integration (the pack is built by hand, see the doc);
  - the FP8 PLE table (#291) and the BF16 embedding (#290), which pair well with it;
  - HIP: the NVFP4 MMQ is guarded off, and the decode kernels use no CUDA-only intrinsics, but I could not build it.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.