Pull requests / #473
#473 native experts: Q5_0 GPU kernel (#222)
closed · @Rhonstin · 0 commentaires · Sur GitHub
BenchmarksNVIDIA / CUDAModels & quants
Description
## Summary
Two related fixes found while getting a community Q4_K_M GGUF (gate/up Q4_K, down Q5_0 on several layers; a 110 B-row Q5_0 PLE table) running on Strata.
### 1. Native expert GPU kernel for Q5_0
`native_expert_supported()`/`is_iq()` already cover Q4_K, Q5_K, Q5_1 and Q8_0 for the native expert GPU path, but not **Q5_0**. A plain `llama-quantize` "_M"/"_S" tier GGUF (not Unsloth's dynamic method) commonly puts Q5_0 on down-projections — sometimes gate/up too — and those layers were refused outright at startup ("this engine has no GPU kernels for"). Fixes #222 for Q5_0 specifically (Q6_K is intentionally still out of scope — see `native_expert_parity_refuses_q6_K`, left untouched).
- `vec_dot_q5_0_q8_1` / `dq_q5_0`: port `vecdotq.cuh`'s integer chain and dequantizer verbatim. Q5_0 is symmetric (one delta, no learned min), so unlike Q5_1 there's no min-term rounding choice to make — the fixed -16 offset is accounted for via the q8_1 block's `ds.y`, exactly as upstream.
- `Fmt<6>`, plus the `STRATA_GU_FMTS`/`STRATA_D_FMTS`/`STRATA_MMVQ_FMTS`/`is_iq`/`iq_row_bytes`/`dq_dispatch` entries, following the existing Q5_1 (type 7) pattern.
- Two entries in the `native_expert_parity` CMake test matrix: `q4_K/q5_0` and `q5_0/q8_0`. Both: 0 failures, GPU dequant bit-exact against `to_float`, gate/up and down rows bit-exact against ggml's own `vec_dot`, width invariance holds.
- `q5_0/q5_0` (Q5_0 in *both* roles) is deliberately **not** added: stacking the same lower-precision format on both roles pushes the GPU path to ~3.6% relative error against the float reference, just over this test's 3% threshold. CPU rel error is the same ~1.25% baseline as `q5_0/q8_0`, so this is compounded quantization noise from using a coarser format twice in a row, not a kernel correctness issue — and I haven't seen a real quantizer that actually picks Q5_0 for both roles of the same expert.
### 2. Lift the stale Q5_0 + `--ple-io direct` refusal (#296)
`PleReader`'s `row_bytes` has been a runtime parameter since the FP8 table (160 B rows) needed it — nothing in the actual read/cache/job-building code (`RowCache`, `Job`/`Use`, the straddle-length formula in `issue()`) is 90-specific. `ngram.cpp`'s "Q5_0 PLE requires --ple-io mmap" refusal predates that generalization (or was just never revisited): FP8's 160 B rows already go through `PleIo::Direct` with no such gate.
- `ple_reader_test --selftest` now runs its full synthetic suite (decode-shaped tickets, page-straddling rows, dedup/edges, bulk, two in-flight tickets, fault injection, keep-alive — both threaded and caller-thread modes) at `row_bytes=110` as well as the production default 90, **before** touching the refusal. Both: OK.
- `ple_reader_test --gguf` against the real OrcaRouter Q4_K_M shard: direct vs mmap over 312 real PLE tickets — bit-identical, and direct is **~8.6x faster** (461 vs 3959 us/token) on this SSD, matching the ~2-2.6 ms/token mmap cost the original "sixteen serial page faults" comment in `gather()` measured for the IQ4_NL table.
## Testing (end to end, not just synthetic)
Verified on the running model (engine 0.1.32, RTX 3090) with real text, math and vision generations, `--ple-io mmap` vs `direct`:
- Same outputs both ways (deterministic sampling) — correctness.
- Decode: ~9-12% faster across three real generations.
- Image-prompt PLE throughput: up to 185 tok/s with `direct` (not practical to measure the equivalent with `mmap`). The full per-token win is smaller than the isolated 8.6x since PLE is one part of a token's cost, not all of it.
## Scope
Four files: `src/kernels/cuda/iq_kernels.cu`, `CMakeLists.txt` (test registration), `src/kernels/ngram.cpp` (the one-line refusal removed), `src/ngram/ple_reader_test.cpp` (parametrized synthetic test, no behavior change to the test itself at `row_bytes=90`). No changes to pack-time classification (`tools/iq_pack.py`'s `native_experts.txt` writer) or to `PleReader`/`DirectFile` themselves — both already worked for arbitrary row sizes.
Sur le site
Liens install, modèles, releases.