Pull requests / #933

#933 CPU trellis kernels (AVX2 / AVX-512) for TQ2_T / TQK6 / TQK7 expert-cache misses (stacked on #928)

closed · draft · @LaurentZuijdwijk · 0 评论 · 在 GitHub 查看

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows

描述

**Stacked on #928.** This PR contains #928's commits plus 3 new ones (the last three in the list); please review it after #928.

On cards that can't hold every expert (16-24 GB), cache misses run on the CPU workers. For the trellis types (TQ2_T / TQK6 / TQK7) those misses went through ggml-cpu's `vec_dot`, one token at a time. This PR adds Strata's own AVX2 and AVX-512 decode-and-dot kernels for those types, wired into `native_gu_rows` / `native_down_rows`.

## What changes
- `src/kernels/cpu/tq_cpu.{hpp,cpp}`: a state table, the scalar reference, activation unpacking and run-time ISA choice. `tq_avx2.cpp` (AVX2 + FMA) and `tq_avx512.cpp` (F/BW/VL/VNNI) hold the kernels.
- **Each 128-weight block is decoded once and reused for every token routed to that expert**, so the extra tokens in an MTP verify window are cheap.
- The Hadamard input rotation is unchanged: once per token per layer.
- **Switches:**
  - `STRATA_NO_TQ_KERNELS=1` goes back to the ggml path;
  - `STRATA_TQ_ISA=scalar|avx2|avx512` caps the ISA;
  - a CPU without AVX2 keeps ggml.
- **Tests:**
  - new CPU-only `tq_cpu_parity` (ctests `tq_cpu_parity` and `tq_cpu_parity_avx2`), with a `--bench` mode;
  - `native_expert_parity`'s synthetic modes also check every ISA against scalar.
- `docs/GYRO.md` has a short note.

## Correctness
- AVX2 and AVX-512 are **bitwise equal** to scalar, for any group of 1-8 tokens. A token's result is identical alone or in a group.
- Against ggml-cpu `vec_dot`: relative error ~1e-7, which is summation order.
- Against `to_float` in double: relative error 1.1e-5, the int16 table rounding that ggml's kernel shares.
- Whole expert, with and without Hadamard, through the engine entry points: 1.2e-7 against the ggml path.
- `native_hadamard_test` passes.
- CPU and HIP (gfx1151) builds pass.

## Speed (core cycles per 128-weight block, Zen 5, one thread, weights from DRAM)

| kernel | 1 token | per token, group of 3 |
|---|---:|---:|
| ggml `vec_dot`, AVX2 build | ~115 | ~115 |
| ggml `vec_dot`, AVX-512 build | ~64 | ~61 |
| **this PR, AVX2** | **~48** | **~22** |
| **this PR, AVX-512** | **~48** | **~22** |
| q2_0 CPU kernel, for reference | 20 / 7.4 (AVX-512 repack) | 12 / 7 |

That is about 2.4x (AVX2 builds) or 1.35x (AVX-512 builds) faster for one token, and 2.8-5x per token in verify windows. Q2_0's simpler format stays cheaper per block on the CPU.

## Measured end to end (2026-10-05)

**RTX 3090 24 GB + AMD EPYC 7K62 (Zen 2, AVX2 only)**, one GPU, `--expert-cache auto`, MTP spec 4, int8 KV, 128k context. Decode tok/s, mean of 6 reps over two rounds with the arm order flipped. Old = `STRATA_NO_TQ_KERNELS=1` (ggml-cpu `vec_dot`); new = this PR (AVX2 here).

| model | cache hit rate | prose | json | code | copy |
|---|---:|---:|---:|---:|---:|
| Gyro-M (TQ2_T), old → new | 93% | 34.2 → 35.6 (+4%) | 50.2 → 53.7 (+7%) | 44.7 → 49.3 (+10%) | 47.8 → 54.2 (+13%) |
| Gyro-S (TQK6/7), old → new | 97% | 38.3 → 38.0 | 55.2 → 57.9 (+5%) | 53.5 → 53.4 | 57.0 → 58.8 (+3%) |
| GSQ-RCO Q2_0 (reference) | 94% | 35.7 | 61.2 | 56.4 | 67.3 |
| GSQ-RCO IQ3_XXS (reference) | 90% | 30.9 | 57.3 | 51.9 | 56.8 |

- **Gyro-M gains +8.7% on average.** While the cache is still warming up (79-93% hits) it gains +10-22%: the gain grows with the miss rate. Gyro-S, at 97% hits, leaves the CPU little to do (+1.8%, within noise on prose and code).
- **Kernel micro-benchmark** (4 pinned cores, same host): 42-48% of ggml-cpu's time per expert; about 19-21% per token in a 3-token verify window.
- **Prefill** is unchanged within 2%; it doesn't use these kernels.
- **Determinism:** greedy output can differ from the old path on some prompts, by float summation order (json identical, one prose prompt diverges after ~500 characters and stays fluent). Across ISAs (scalar/AVX2/AVX-512) the new kernels are bitwise equal.
- Host load average on the rented box was 12-31 between runs, so treat ±3% as noise.

**RTX 4090 + Ryzen 9 7950X (Zen 4, AVX-512)**, Gyro-S, 97% hit rate: old and new are equal within noise. At that hit rate the CPU does little of the expert work. On the same box Strata runs Gyro-S at 127.5 tok/s mean decode at 119k context (reasoning / bug fix / edit 113 / 128 / 142), with 4,177 tok/s cold prefill on 118.7k tokens.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_013YKZhadEzjDiTwDc4tzi2g

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。