Pull requests / #651
#651 ple: read Q8_0 and Q5_1 n-gram tables
closed · @anon761 · 0 comentários · No GitHub
BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants
Descrição
## What Several ordinary GGUFs of Qwen3.8-Flash-Next ship `per_layer_token_embd.weight` in a format the PLE reader refuses at start: | file | PLE table | | --- | --- | | Unsloth `UD-Q6_K_XL` | Q8_0 (170-byte rows, ~54 GB) | | Swift-1.5 `Q4_K_L` | Q8_0 | | a community `Q5_K_M` | Q5_1 (120-byte rows) | ``` strata generate: per_layer_token_embd.weight is Q5_1, not IQ4_NL, Q5_0 or FP8 (I8) ``` The rest of these files already has a path in the engine (Q5_K/Q6_K/Q5_1/Q8_0 experts, the K-quant projections, `--compat-bf16`). For the Q5_K_M file the table was the only thing keeping it from starting (end to end below); the Q8_0 tables are covered by the parity test on the real files. - `dequantize_q5_1` in `strata/artifact/dequant.hpp` (ggml's `dequantize_row_q5_1`: unsigned 5-bit codes times `d` plus the block minimum `m`); Q8_0 rows use the existing `dequantize_q8_0` - `PleTable` decodes Q5_1 / Q8_0 rows on both readers. The direct reader already takes the row size at runtime (since the FP8 table), so no reader change was needed; `PLE_ROW_BYTES_MAX` grows to 170 - `--ple-io ram`: before `mlock`, 16 threads fault the table in. `mlock` alone brings a cold file in one page at a time from one thread; on a ZFS pool that read the 54 GB Q8_0 table at ~0.47 GiB/s (about two minutes), 16 threads at ~2 GiB/s (measured from a cold cache before the rebase onto current `main`; the container this was re-tested in cannot drop the file cache). The fallback when `mlock` fails keeps the pages faulted in instead of touching them a second time. ## Tests - `dequant_q5_1_test` (new, CTest): 4096 synthetic blocks incl. the edge codes 0 and 31 and a negative minimum against ggml's `to_float`, bit-exact. No model, no GPU. - `dequant_bf16_test` knows Q5_1. - `ple_q5_parity` now takes a Q5_0, Q5_1 or Q8_0 table and checks **both** the mapped and the direct reader (single rows and a 16-row batch) against ggml's `to_float` on the real file: ``` Q5_1 PLE, mapped reader: 320001536 rows, max_abs 0.000e+00 PASS (Q5_K_M, shard 1) Q5_1 PLE, direct reader: 320001536 rows, max_abs 0.000e+00 PASS Q8_0 PLE, mapped reader: 320001536 rows, max_abs 0.000e+00 PASS (Swift-1.5 Q4_K_L, shard 2) Q8_0 PLE, direct reader: 320001536 rows, max_abs 0.000e+00 PASS Q8_0 PLE, mapped reader: 320001536 rows, max_abs 0.000e+00 PASS (UD-Q6_K_XL, shard 3) Q8_0 PLE, direct reader: 320001536 rows, max_abs 0.000e+00 PASS ``` ## End to end The Q5_K_M file (Q5_1 table) on 2x RTX 3090, layer split auto, 262K context, `--kv int8 --pcie-frac 0`; `main` refuses it at start: | `--ple-io` | prompt 2K | prompt 8K | decode | | --- | ---: | ---: | ---: | | `ram` (table locked in 3.2 s, warm file cache) | 705 tok/s | 2116 tok/s | 85-87 tok/s | | `direct` (default) | 703 tok/s | 2117 tok/s | 65-74 tok/s |
No site
Links install, modelos, releases.