Pull requests / #1401

#1401 V100 (sm_70): prompt experts on FP16 tensor cores (+10% prompt) and three existing decode kernels as defaults (-6.8% GPU per window)

open · @rewin123 · 0 Kommentare · Auf GitHub

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

Beschreibung

## Title
V100 (sm_70): prompt experts on FP16 tensor cores (+10% prompt speed) and three existing decode kernels as Volta's defaults (-6.8% GPU work per window)

## Summary
Two parts, both picked only on sm_70 (every other card runs what it ran), each with a switch back to the old path.

**Prompt (+9.6% / +10.3%).** MMQ's integer products are `dp4a` on Volta. `gemm_iq_f16_grouped` runs an MMQ expert group per launch on the FP16 tensor cores (WMMA, FP32 sums): each weight superblock dequantized once into shared memory with `dq_dispatch` (the FP16 path's formulas), FP16 activations streamed through, the next tile prefetched. It also takes the IQ1_M layers through MMQ's grouping (they were on the per-expert FP16 path). V100-SXM2, setup's IQ2_XS and config, 3 alternating rounds: 26,359-token prompt 1,467 -> 1,608 tok/s, 6,345-token prompt 1,390 -> 1,534 tok/s (rounds within 1%). `STRATA_PF_WMMA=0`: MMQ.

**Decode (bitwise).** Three kernels Strata already has for other cards: gfx906's expert mode 8 for the routed experts in VRAM, `native_mmvq_il` with a Volta rows table, gfx906's latency-hidden hyper-connection up. GPU work per verify window 22.80 -> 21.24 ms (engine stage table, 3 rounds); decode tok/s +3-11% against the same-binary control on this rig, whose CPU pool makes decode noisy. `STRATA_EXP_MODE=0`, `STRATA_MMVQ_IL=0`, `STRATA_GR_FAST=0`.

Numbers, method, raw data: `bench/results/2026-10-07-v100-prompt-experts/` and `bench/results/2026-10-07-v100-decode-kernels/`.

## What changed
- `src/kernels/cuda/iq_kernels.cu`: `gemm_iq_f16_grouped` (CUDA; a stub on HIP); the AMD grouped-expert layouts (`STRATA_EXP_MODE`) build for CUDA, sm_70 defaults to mode 8, IQ2_XXS gate/up joins the shared-memory kernels on CUDA, a strided call (the window's PCIe call) and a format without an LDS kernel keep the CUDA layout. gfx906 unchanged.
- `src/prefill/prefill.cpp`: `pf_wmma()` (sm_70 default), IQ1_M layers in MMQ's grouping for it, the FP16 activation gather instead of the q8_1 quantization on those layers, the group's two tile launches + an FP16 SwiGLU (`src/prefill/kernels.cu: swiglu_split_f16`).
- `src/kernels/cuda/native_mmvq.cu`: the interleaved 2-4 column kernel on sm_70 with its own rows table (read off `mmvq_il_parity --bench` on the V100 by the existing table's rule).
- `src/kernels/cuda/fused_gr.cu`: gfx906's fast norm/up build for CUDA, default on sm_70; `gr_up_fast_kernel` now writes the QFUSE q8_1 tail like `gr_up_multi_kernel` (this also fixes gfx906 with `STRATA_QFUSE=1`).
- Tests/tools/docs: `src/kernels/pf_wmma_parity.cpp` (+ ctest), `tools/ab_engine.py`, `docs/NVIDIA_V100.md`, the two bench result folders.

## Extra Notes
- Prompt accuracy: `pf_wmma_parity` against a double reference over the same FP16 weights: max error ~5e-6 of the output scale. Teacher-forced (`STRATA_LOGPOS`, a continuation after a 26K / 12.6K prompt): against MMQ, KL 0.0083 (prose) and 0.031 (code) - the same size as MMQ against itself with another chunk size (0.0084-0.0177 prose, 0.025 code). MMQ is not a reference (q8_1 activations); a full-precision reference run was not made.
- Decode changes are bitwise the kernels they replace (`native_expert_bench` on all 48 layers, `native_grouped_parity`, `mmvq_il_parity`, `mmvq_multi_parity`, `fused_gr_bench`, `gr_multi_parity`, `gr_parity`).
- Full `ctest` on the V100: all pass except 3 that fail on this rig only (no AVX512-VNNI, `mlock` in a container, the Q2_0 shard `ple_parity` wants); `prefill_fused_*` skip below sm_80.
- Not measured: a PCIe V100, other packs (Q2_0 / IQ3_S / UD-IQ4_XS take the same code; the prompt kernel needs a Q2_0 down product, so UD-IQ4_XS's IQ4_NL down keeps MMQ), HIP builds (unchanged by construction, not built here).
- The history has one change reverted in place (the FP16 path for 16K+ chunks, superseded by the tensor-core kernel); squash-merge reads cleanest.
- Pointed at by 1Cat-vLLM's SM70 GGUF kernels (shared-memory lattice codebooks, tensor-core expert GEMMs); no code was taken from it.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01AUz3LpAVcsKzRLDoJSkZP4

Mehr auf der Site

Links zu Install, Modellen, Releases.