Pull requests / #292
#292 NVFP4 routed experts: converter, pack, decode, prompt path (W4A8, W4A4 on Blackwell), AVX-512 CPU rows
closed · @sergqwer · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows
描述
## Summary
A ModelOpt NVFP4 checkpoint of Qwen3.8-Flash-Next runs on the native path: routed experts in NVFP4 (group 16), the
rest in the formats the native path already reads. `docs/NVFP4.md` has the whole design and every measurement.
These changes come from an NVFP4 fork of Strata (sergqwer/strata-nvfp4), ported onto current main one topic per PR.
**Pipeline.**
- **Converter** (`tools/nvfp4_convert.py`): llama.cpp's own converter (the pinned commit) with the engine's type
policy. The experts are repacked without loss; `tools/nvfp4_verify.py` and `nvfp4_verify_gguf.py` compare them with
the checkpoint bit for bit.
- **Pack** (`tools/iq_pack.py`): each NVFP4 blob gets a 16-byte tail `{s_gate, s_up, s_down, 0}`, the per-expert
`weight_scale_2`. Every consumer that already moves whole blobs (arena, device cache, prefill staging, a second
GPU) then moves the scales with the weights. `experts.bin` is written by itself, because it is the only source with
the tails.
- **Decode**: NVFP4 in the native expert kernels (q8_1 activations, UE4M3/E2M1 decoded in software, bit-identical to
ggml). Each global scale multiplies its own projection's FP32 output: folding `s_down` into up put the hidden into
FP16's subnormals (2-12% expert error instead of 1.1%).
- **CPU pool**: AVX-512 NVFP4 rows, 1.8-3.7x faster than ggml-cpu's (identical arithmetic, another summation order);
ggml-cpu without AVX-512.
- **Prompt path**: llama.cpp MMQ. The default is **W4A8**: the int8 MMQ, compiled for sm_120 with Blackwell's FP4 MMA
hidden (`mmq_nvfp4_w4a8.cu`). **W4A4** (FP4 MMA, sm_120a) and **FP16** are opt-in through `STRATA_PREFILL_NVFP4`.
- **PCIe share**: 0.25 of the missed experts for NVFP4 packs, not 0.55. Their blobs are 2.76 MB, and at 0.55 the copy
kernel was 40% of GPU time.
- **Build**: CMake turns 120 into 120a (as ggml's own CMake does), adds the NVFP4 MMQ instance and the parity tools,
and guards the CUDA-only parts for HIP.
## Measured
RTX 5090, Ryzen 9 9950X3D, 128 GB, Windows 11; this branch on main, 64K context, int8 K/V, `--vram-reserve-mib 1500`:
| | NVFP4 pack | IQ2_XS, same engine |
| --- | ---: | ---: |
| experts | 63.3 GiB | 33.0 GiB |
| expert cache slots | ~8,350 | ~17,300 |
| decode, 400-token answers | 114-116 tok/s | 132-139 tok/s |
| a 32K prompt | ~3,900 tok/s | ~5,600 tok/s |
Precision of the prompt path, first-token KL against FP16, 8 prompts:
| mode | KL mean | top-1 |
| --- | ---: | ---: |
| W4A8 (default) | 0.0018 | 8/8 |
| W4A4 | 0.0080 | 7/8 |
| noise floor | 0.00023 | 8/8 |
W4A8 reads prompts at 2,503 tok/s, W4A4 at 2,904 and FP16 at 1,905 (4-8K prompts, 262K context).
## Tests
- `nvfp4_expert_gpu_parity`: the decode kernels against ggml's reference.
- `nvfp4_avx512_parity`: the CPU rows against ggml-cpu.
- `mmq_nvfp4_parity`: the prompt path's products against FP64.
- The verify scripts: the conversion, bit for bit.
- sm_75/86/89 builds (emulated on the RTX 5090): the parity tools give results identical to sm_120a.
## Tested with, and what is not in this PR
- **Checkpoints.** Run end to end with `jpezzulli/OrcaRouter-Qwen3.8-Flash-Next-Uncensored-ModelOpt-NVFP4`.
`nvidia/Qwen3.8-Flash-Next-NVFP4` quantizes its routed experts the same way (by its `hf_quant_config.json`), so it
takes the same path. I have not run it.
- **Not included:**
- setup integration (the pack is built by hand, see the doc);
- the FP8 PLE table (#291) and the BF16 embedding (#290), which pair well with it;
- HIP: the NVFP4 MMQ is guarded off, and the decode kernels use no CUDA-only intrinsics, but I could not build it.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。