Pull requests / #353
#353 NVFP4 routed experts on 0.1.35 (experimental): 120a only for the W4A4 unit (rebase of #292)
closed · @sergqwer · 0 commentaires · Sur GitHub
BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux
Description
The rebase asked for in #292, for 0.1.33 as experimental. One commit on 0.1.31 (9259cad). ## What changed against #292 - **Rebased on 0.1.31.** - NVFP4 is type 40 in the per-role format lists (`STRATA_GU_FMTS` / `STRATA_D_FMTS` / `STRATA_MMVQ_FMTS`), next to Q4_K/Q5_K/Q5_1/Q8_0. It goes through the same `launch_*` dispatch. - 0.1.31's `dq_q8_0` serves the Q8_0 token embedding as well; #292 had its own copy, which is gone. - `tools/iq_pack.py`: `expert_layout()` gives NVFP4 blobs their 16-byte scale tail. An NVFP4 pack always writes `experts.bin`, because the GGUF has no tails and so the in-place GGUF read does not apply. - **120a is limited to the W4A4 unit.** CMake no longer rewrites `CMAKE_CUDA_ARCHITECTURES`. - `src/prefill/mmq_nvfp4_w4a4.cu` is the only unit built with `12x` -> `12xa`, in its own static library `strata_mmq_w4a4`. It holds the FP4 x FP4 kernel, its FP4 activation quantizer and a probe, with their symbols renamed so they cannot fold with the plain-arch copies. - Everything else, the default W4A8 path included, keeps the architectures as given. `strata.exe` from `-DCMAKE_CUDA_ARCHITECTURES=120` holds 55 sm_120 cubins and 1 sm_120a. - When the probe has no image for the running card (another architecture, or sm_121 against a 120a build), `STRATA_PREFILL_NVFP4=w4a4` falls back to W4A8 with a note. - The stock `mmq-instance-nvfp4.cu` is no longer built. - HIP builds never reference the W4A4 unit. ## Checked on this rebase (RTX 5090, Windows) - **Parity tools:** - `nvfp4_expert_gpu_parity`: decode experts 1.09-1.16%. - `nvfp4_avx512_parity`: OK. - `mmq_nvfp4_parity`: W4A8 0.53%, W4A4 8.6%. These are bit for bit the numbers of the old 120a-everywhere build. - **IQ2_XS:** 128 greedy tokens identical to 0.1.31 main and to 0.1.29. - **NVFP4 pack against the 0.1.29-based #292 build, same `--expert-cache` size:** identical tokens on a 24-token prompt and on an 8K prompt, both in W4A8 and in W4A4. Prefill on 8K is 4,231 tok/s in W4A8 and 4,583 in W4A4; decode is the same. - **`--expert-cache auto` gives slightly different slot counts on 0.1.31** (8,498 against 8,472 on 0.1.29). That changes which misses the CPU pool computes. For NVFP4 the AVX-512 rows equal the GPU's only to a relative tolerance (other float summation order), not bit for bit. The identity gate should therefore compare NVFP4 at a fixed `--expert-cache N`. - **Run-to-run variation:** once in five runs, an NVFP4 run diverged from the others late in a low-confidence stretch. This is presumably the same effect: the CPU/GPU split of the misses depends on timing. The IQ formats are unaffected, since their CPU rows are bit-exact. Not run here: Linux, AMD (HIP), sm_121. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
Sur le site
Liens install, modèles, releases.