Pull requests / #860
#860 Unsloth UD-Q6_K_XL: Q8_0 PLE rows and Q6_K gate/up experts
closed · @DaRealDaHoodie · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentationLinux
描述
Unsloth's UD-Q6_K_XL of Qwen3.8-Flash-Next (revision `38bb39e`) does not load on main. Its PLE table is Q8_0 (170-byte rows, 54.4 GB) and its routed gate/up experts are Q6_K. Main reads the table as IQ4_NL, Q5_0 or FP8, and `native_expert_supported` refuses Q6_K gate/up with "no GPU kernels for". This branch teaches the engine those two formats and documents the manual pack. Setup does not offer the model. ## What changed - `PleTable` accepts `per_layer_token_embd.weight` as Q8_0. The row reader was already parameterized by bytes per row; this adds the type check, a `dequantize_q8_0` loop (5 blocks of 34 bytes) and `PLE_ROW_BYTES_MAX` of 170. Gather still writes floats, so no GPU kernel changed. v0.1.39's PLE prefetch (up to 16 rows) stays; with `--ple-io ram` that path is idle. - Grouped expert kernels take gate/up Q6_K (`Fmt<14>`, llama.cpp's `vec_dot_q6_K_q8_1`). Per-token MMVQ already had Q6_K. `STRATA_MMQ_KQUANTS=ON` also builds the Q6_K MMQ prompt instance. - The pack on this file is gate/up Q6_K and down Q8_0 on 47 of 48 layers. Layer 2 is Q8_0 for both. Down Q6_K is not in this file. - `native_expert_parity --synthetic q6_K/q8_0` is registered as a normal test. It used to be `WILL_FAIL`, from when that pair had no GPU kernel. ## Measured Linux, RTX PRO 4500 Blackwell (compute capability 12.0) plus a second GPU, Threadripper PRO 5975WX (AVX2, no AVX-512), 30 expert-pool workers, 499 GB of RAM. CUDA archs 86;120, `STRATA_MMQ_KQUANTS=ON`. Experts resident in RAM (101.75 GiB), `--expert-cache` 5465 slots / 22.63 GiB, CUDA1 1800 experts / 7.48 GiB, `--kv-resident 32768`, `--ple-io ram`, `--pcie-frac 0.20`, `--spec 4`. Engine 0.1.39, one review, 2026-10-04: - Cold read of 14,915 tokens at 3,176 tok/s. - 14 turns, 5,799 tokens in 74.1 s, 78.3 tok/s, ending at 44,160 tokens of context. - Drafts accepted 3,485 of 4,406 (79.1%). Same machine, engine 0.1.38, before this branch was rebased onto 0.1.39: a cold read of 249,193 tokens at 2,802 tok/s, then 1,079 tokens at 69.0 tok/s. `ple_reader_test --selftest` covers 170-byte rows, and `ple_q8_parity` matches ggml's Q8_0 dequantizer bit for bit on the synthetic fixture (`max_abs 0`). ## Not measured Quality against llama.cpp on this file. The three Q6_K parity programs (`native_grouped_parity`, `prefill_mmq_kquant_test`, `native_expert_parity --synthetic q6_K/q8_0`) are built and have not been run: the GPU was serving. The page stays experimental until those exist. `docs/UNSLOTH_Q6.md` is the manual workflow.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。