Pull requests / #928

#928 Gyro (TQ2_T / TQK6 / TQK7 + Hadamard rotation): native support for the agentionai Qwen3.8-Flash-Next quants

closed · @LaurentZuijdwijk · 0 评论 · 在 GitHub 查看

BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

描述

Native support for the agentionai **Gyro** quants of Qwen3.8-Flash-Next (`agentionai/Qwen3.8-Flash-Next-Gyro-GGUF`), so Strata can serve them unchanged next to the GSQ-RCO files.

## What's in it
- **Trellis expert types TQ2_T / TQK6 / TQK7** (128-weight blocks, one fp16 scale; 2.125 / 1.625 / 1.875 bits per weight), read straight from the GGUF, with decode kernels on CUDA, HIP and the CPU path. On CUDA the codebook is decoded as int8 and the dot products use dp4a.
- **Block-128 Hadamard activation rotation** for experts stored in a rotated basis (`prism.hadamard.*` metadata): rotation on the GPU decode and prompt paths, the CPU workers, the remote/peer expert paths and the stage helper. Fused prompt kernels skip rotated layers and use the separate swiglu, rotate, quantize passes.
- **Readers** for a Q6_K token embedding and a Q8_0 per-layer (n-gram) table, as the Gyro files use them.
- **Multi-GPU:** layer split with the prompt helper, and decode with `--peer-device`. Prompt processing on the peer card needs P2P and is untested.
- **`tools/gyro_setup.sh`** (build, download, pack, MTP draft head, config with `--prefill auto`) and **`docs/GYRO.md`**.

Strata needs a GGML with the trellis types: `STRATA_GGML_DIR` should point at https://github.com/agentionai/llama.cpp (`main`).

## Tested
- **CUDA builds:** sm_120 (RTX 5090), sm_89 (2× RTX 4090), sm_86 (RTX 3090 / A6000).
- **HIP builds:** gfx1151 (Strix Halo, ROCm 7.2.1).
- **`native_expert_parity`:**
  - `--synthetic` and `--synthetic-hadamard` for tqk6/tqk7 and tq2_t pass with 0 failures on CUDA and gfx1151;
  - dequant is bit-exact against ggml's `to_float`;
  - the GPU rotation equals the CPU's bitwise;
  - `--load-check` passes with 0 failures for Gyro-S and Gyro-M packs.
- **End to end:**
  - Gyro-S on RTX 5090, RTX 4090, RTX 3090 and Strix Halo; Gyro-M on RTX 5090 and RTX 3090;
  - output reads correctly on prose, code and JSON.
- **Two GPUs (2× RTX 4090, no P2P):**
  - the layer split with `STRATA_PREFILL_HELP=1` gives identical output over repeated runs;
  - `--peer-device 1` and `--expert-cache-device1 auto` decode match single-GPU.

## Numbers
Quality is held-out KLD against the Q8_0 source, `-c 2048`, 60 chunks; GPU-resident size excludes the n-gram table.

| file | GPU-resident | KLD | top-1 |
|---|---:|---:|---:|
| Gyro-S | 27.64 GiB | 0.435 | 74.1 % |
| Gyro-M | 34.98 GiB | 0.306 | 78.1 % |
| GSQ-RCO Q2_0 | 35.03 GiB | 0.468 | 73.9 % |
| GSQ-RCO IQ3_XXS | 43.80 GiB | 0.311 | 77.7 % |

Strata decode speed (tok/s, prose / JSON / code):

| card | Gyro-S | Gyro-M |
|---|---|---|
| RTX 5090, all experts cached | 140 / 221 / 209 | 114 / 171 / 153 |
| RTX 3090, 24 GB with the expert cache (older server CPU) | 32 / 45 / 42 | 28 / 39 / 34 |
| Strix Halo (HIP) | ~21 / 31 / 30 | |

On 24 GB cards, cache misses run the trellis decode on the CPU. That costs more than Q2_0's, and a SIMD CPU path is the next step.

## Not yet verified
- Prompt-path peer offload with these files (needs P2P/NVLink).
- No speed tuning on HIP yet.
- Any interaction with your upcoming Strix changes.

Happy to rebase onto whatever you'd like and to split this into smaller PRs if that's easier to review.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_013YKZhadEzjDiTwDc4tzi2g

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。