贡献 / #1424

#1424 Pascal GP100: a Q8_0 decode GEMV (STRATA_Q8_SM60=1, opt-in) and IQ3_XXS / IQ3_S tables in shared memory

open · @ruibeikaa · 0 评论 · 去 GitHub 看

BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows

说明

## Title
Pascal GP100: a Q8_0 decode GEMV (`STRATA_Q8_SM60=1`, opt-in) and IQ3_XXS / IQ3_S tables in shared memory

## Summary
Two sm_60 changes for Pascal GP100 (the chip from #124), measured on a PH402 SKU 200: two boards, four GP100 dies of
48 SMs and 32 GB HBM2 each.

GP100 is the one Pascal chip without `__dp4a`, and on a GP100 die the decode window's dense projections read their
weights at 40-80 GB/s of ~730: the K- and i-quant blocks are decoded one int at a time. With `STRATA_Q8_SM60=1` and a
shard whose dense projections and head are Q8_0, decode goes from **28.3 to 38.9 tok/s (+37%)** on Flash-Next
IQ3_XXS; prompt reads are unchanged. Everything is off by default except the IQ3_XXS and IQ3_S tables, which are
bitwise the same output and on by default for compute capability 6.0 (GP100) only; the IQ3_S ones add **+3.2%
decode** on Swift 1.5 IQ3_XXS.

| Flash-Next IQ3_XXS, 32K, decode tok/s | ~10.7K-token document | the same, cached | short reasoning | code | mean |
| --- | ---: | ---: | ---: | ---: | ---: |
| GSQ-RCO shard, exact kernels | 28.8 | 29.4 | 28.4 | 26.8 | 28.3 |
| Q8_0 dense shard, `STRATA_Q8_PACKED=1` | 27.5 | 31.0 | 34.4 | 31.8 | 31.1 |
| Q8_0 dense shard, `STRATA_Q8_SM60=1` | 37.5 | 38.4 | 41.2 | 38.5 | 38.9 |

GPU time per verify window (`STRATA_VERIFY_PROFILE`): dense projections 19.2 -> 11.2 ms, head 4.25 -> 2.1 ms; the
draft policy then picks longer windows (2.10 -> 2.31 tokens per window). Prompt reads: 437-438 tok/s at 10.7K in
every row.

Conditions: application clocks locked at 1050 MHz, Windows 11, driver 581.80 (TCC), CUDA 12.9, engine 0.1.40.2 (the
branch is rebased on 0.1.40.3, which touches none of these files), `--layer-split 12,24,36 --prefill 2048`, greedy,
256 new tokens, medians of 2-3 interleaved engine starts, the engine kept on its own logical CPUs (Extra Notes).

## What changed
- `src/kernels/cuda/q8_sm60.cuh` (new): the GP100 Q8_0 GEMV. A warp (half / quarter warp for rows under 2048 / 1024
  wide) per row; 16-byte weight loads through the read-only cache; the window's q8_1 activations copied once per
  workgroup into shared memory as an int8 plane plus float scales, so every activation read is a 16-byte shared load;
  `STRATA_DP4A` (vmad on sm_60); a shuffle sum with no barrier per row; persistent workgroups, 4 per SM
  (`STRATA_Q8_SM60_BPS`). Shapes needing more than 48 KiB of shared memory (a 6144-wide row with 7-8 columns) are
  declined and keep the usual kernel. Also a device-side GGUF-block -> planes repack for the head.
- `src/kernels/cuda/native_mmvq.cu`, `include/strata/kernels/native_mmvq.hpp`: `STRATA_Q8_SM60=1` makes every Q8_0
  matrix eligible for the packed planes and dispatches registered matrices to the new kernel (any ncols it accepts,
  independent of the exact-layout requirement); `native_q8_0_sm60_pack` / `native_q8_0_sm60_release` own the head's
  packed copy. Not compiled into HIP builds. Without the variable nothing changes.
- `src/core/native_head.cpp`: a Q8_0 head is packed when `STRATA_Q8_SM60=1` (the MTP draft head reads the same matrix).
- `src/kernels/cuda/iq_kernels.cu`: IQ3_XXS and IQ3_S join the types whose codebook the grouped expert kernels stage
  in shared memory, each with sign-mask tables applied as `(g ^ mask) + (mask & 0x01010101)` (two 128-word tables for
  IQ3_XXS's 7-bit sign fields, two 256-word tables for IQ3_S's sign bytes; no codebook byte is 0, so a negated byte
  never carries). IQ3_XXS's are ported from
  [shinbunbun/llama-cpp-p100-patches](https://github.com/shinbunbun/llama-cpp-p100-patches) (MIT), patches 29
  `mmvq-iq3xxs-grid-smem` and 30 `mmvq-ksigns-smem`; IQ3_S's are the same idea - IQ3_S was the one gate/up format
  whose codebook these kernels still read from global memory. Bitwise the same output (`iq_multi_parity` and
  `native_grouped_parity` on a GP100 die; identical greedy replies, and the same drafts accepted in the A/B below).
  IQ3_XXS's alone: within noise on GP100 (+0-4%). IQ3_S's: a verify window's layer of real IQ3_S experts (20 experts,
  24 entries) on one die 517 -> 393 us with a Q2_0 down projection and 492 -> 369 us with IQ4_NL; Swift 1.5 IQ3_XXS
  (13 of its 48 layers IQ3_S), 1M context, otherwise the conditions below, decode tok/s:

  | Swift 1.5 IQ3_XXS, `STRATA_Q8_SM60=1` | ~10.7K-token document | the same, cached | short reasoning | code | mean |
  | --- | ---: | ---: | ---: | ---: | ---: |
  | IQ3_XXS tables | 36.0 | 36.3 | 43.1 | 41.6 | 39.3 |
  | IQ3_XXS + IQ3_S tables | 37.2 | 37.4 | 44.5 | 43.0 | 40.5 |

  Medians of three interleaved engine starts per build; on every prompt each of the three runs with both tables was
  faster than each run without. The staging only pays where these kernels are decode-bound: on an RTX 3090 Ti (sm_86),
  forced on, the same IQ3_S layers took 3-13% longer (71-76 -> 75-84 us; there they already read at 490-665 GB/s) and
  IQ3_XXS was unchanged. So the default is compute capability 6.0 only (GP100 has no `__dp4a`; 6.1 cards do and are
  unmeasured); `STRATA_IQ_STAGE_TABLES=1|0` forces it on or off on any card.
- `tools/q8_dense_gguf.py` (new): the GSQ-RCO packs have no Q8_0 dense matrices, so this writes a copy of the model
  whose natively served projections and head are Q8_0 (Flash-Next IQ3_XXS: 301 tensors, +3.5 GiB, requantization
  error at most 0.4% of a tensor's |w|max). Every shard that holds such a matrix is converted - Swift 1.5 keeps 78 of
  them in shard 1 and 223 in shard 2 - and the others are hard-linked (`--copy` copies). The tensor-info block keeps
  its byte length, so every expert keeps its absolute offset and the existing native pack is reused.
- `tools/q8_sm60_check.cu` (new, built by hand like `sm60_dp4a_check.cu`): 57 shape x column cases (the Flash-Next
  Q8_0 shapes, 1-8 columns) against a CPU double reference, max relative error < 1e-5, then times them.
- `docs/OLDER_GPUS.md`: a "Pascal GP100: a Q8_0 decode kernel (opt-in)" section with the table, the conditions and how
  to use it; the P100 row of the support matrix points to it.

## Extra Notes
- Output: `STRATA_Q8_SM60` is not bitwise the exact kernels (Q8_0 requantization of the dense matrices and another
  summation order). Greedy replies to five prompts (arithmetic, logic, Chinese, code, a question on a 10K-token
  document) gave the same answers, with wording differences after the first sentence or two.
- VRAM: the packed copies sit beside the GGUF layout the prompt path reads (Flash-Next: 2.9 GiB of dense projections
  over the stages, 0.6 GiB for the head).
- Two things mattered as much on this card, and are written up in the docs section: its default application clock is
  759 MHz and a layer split leaves each die idle most of a window, so the dies decode below 1050 MHz unless locked
  (`nvidia-smi -ac 715,1050`; 25.0 -> 29.1 tok/s at 10.7K context); and a second engine on the same PC that pins its
  threads to the same logical CPUs halved this one's decode while it ran - keeping them on different logical CPUs
  (SMT siblings) avoided it.
- Not measured on compute capability 6.1 (P40, P4, GTX 10 series), which has `__dp4a`. Both changes build for sm_61,
  and where `__dp4a` exists the Q8_0 kernel's dot is the hardware instruction (`tools/q8_sm60_check.cu` passes on an
  RTX 3090 Ti). Numbers from a 6.1 card with `STRATA_Q8_SM60=1` (and a Q8_0 dense shard) or with
  `STRATA_IQ_STAGE_TABLES=1` would show whether either helps there.
- Built with CUDA 12.9 and MSVC 2022 (`-DSTRATA_EXPERIMENTAL_SM60=ON -DCMAKE_CUDA_ARCHITECTURES=60`); also for sm_86,
  where `iq_multi_parity` and `native_grouped_parity` pass with the tables on and off. Not built for HIP; the new kernel
  only runs with the variable set, and HIP builds leave it out.



本站相关内容

相关页面的快捷入口。