Pull requests / #353

#353 NVFP4 routed experts on 0.1.35 (experimental): 120a only for the W4A4 unit (rebase of #292)

closed · @sergqwer · 0 comments · View on GitHub

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

Description

The rebase asked for in #292, for 0.1.33 as experimental. One commit on 0.1.31 (9259cad).

## What changed against #292

- **Rebased on 0.1.31.**
  - NVFP4 is type 40 in the per-role format lists (`STRATA_GU_FMTS` / `STRATA_D_FMTS` / `STRATA_MMVQ_FMTS`), next to Q4_K/Q5_K/Q5_1/Q8_0. It goes through the same `launch_*` dispatch.
  - 0.1.31's `dq_q8_0` serves the Q8_0 token embedding as well; #292 had its own copy, which is gone.
  - `tools/iq_pack.py`: `expert_layout()` gives NVFP4 blobs their 16-byte scale tail. An NVFP4 pack always writes `experts.bin`, because the GGUF has no tails and so the in-place GGUF read does not apply.
- **120a is limited to the W4A4 unit.** CMake no longer rewrites `CMAKE_CUDA_ARCHITECTURES`.
  - `src/prefill/mmq_nvfp4_w4a4.cu` is the only unit built with `12x` -> `12xa`, in its own static library `strata_mmq_w4a4`. It holds the FP4 x FP4 kernel, its FP4 activation quantizer and a probe, with their symbols renamed so they cannot fold with the plain-arch copies.
  - Everything else, the default W4A8 path included, keeps the architectures as given. `strata.exe` from `-DCMAKE_CUDA_ARCHITECTURES=120` holds 55 sm_120 cubins and 1 sm_120a.
  - When the probe has no image for the running card (another architecture, or sm_121 against a 120a build), `STRATA_PREFILL_NVFP4=w4a4` falls back to W4A8 with a note.
  - The stock `mmq-instance-nvfp4.cu` is no longer built.
  - HIP builds never reference the W4A4 unit.

## Checked on this rebase (RTX 5090, Windows)

- **Parity tools:**
  - `nvfp4_expert_gpu_parity`: decode experts 1.09-1.16%.
  - `nvfp4_avx512_parity`: OK.
  - `mmq_nvfp4_parity`: W4A8 0.53%, W4A4 8.6%. These are bit for bit the numbers of the old 120a-everywhere build.
- **IQ2_XS:** 128 greedy tokens identical to 0.1.31 main and to 0.1.29.
- **NVFP4 pack against the 0.1.29-based #292 build, same `--expert-cache` size:** identical tokens on a 24-token prompt and on an 8K prompt, both in W4A8 and in W4A4. Prefill on 8K is 4,231 tok/s in W4A8 and 4,583 in W4A4; decode is the same.
- **`--expert-cache auto` gives slightly different slot counts on 0.1.31** (8,498 against 8,472 on 0.1.29). That changes which misses the CPU pool computes. For NVFP4 the AVX-512 rows equal the GPU's only to a relative tolerance (other float summation order), not bit for bit. The identity gate should therefore compare NVFP4 at a fixed `--expert-cache N`.
- **Run-to-run variation:** once in five runs, an NVFP4 run diverged from the others late in a low-confidence stretch. This is presumably the same effect: the CPU/GPU split of the misses depends on timing. The IQ formats are unaffected, since their CPU rows are bit-exact.

Not run here: Linux, AMD (HIP), sm_121.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.