Pull requests / #464
#464 ple: read the n-gram table at full BF16 precision
closed · @constantindjonkam · 0 comentários · No GitHub
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quants
Descrição
## What Stop reading the PLE n-gram table from a lossy quantization of it. The table reaches the engine as **IQ4_NL** - ISTA-DASLab's shard 2 - which is about **8% off per row** against the checkpoint's own values. The checkpoint ships the table in **F8_E4M3** with a single `weight_scale`, at 2.66%, and in **BF16**, which is the source of record. This PR adds a packer (`tools/ple_table_pack.py`, renamed from `ple_fp8_pack.py`) that copies whichever form the checkpoint has, byte for byte - nothing is decoded, rounded or re-quantized - and a reader for the BF16 form. Thank you for the measurement on 0.1.37 / RTX 5090, and for building the BF16 table with this PR's tool against OrcaRouter's checkpoint. Two things from it that this PR now leans on: - **FP8 is already a good table.** Median KL over whole answers 0.00087 against BF16, with the same BF16 table at another cache size giving 0.00080 - i.e. below run-to-run noise - for 51 GB against 102 GB. That is the right call for your fork and the reason this PR leads with the packer rather than with BF16. - **Reading the shipped FP8 bytes is the faithful operation, and a cast would not be.** Your observation that only 6% of the checkpoint's BF16 values have the low 4 mantissa bits zero shows the FP8 table is a real re-quantization, not an upcast - which is exactly why the packer copies bytes instead of casting. ## What the review did not measure, and what we measured instead Your comparison is FP8 against BF16. The table a stock install actually runs is **IQ4_NL**, three times further out again, so that is the gap worth closing. First-window logits, same engine, same prompt, fixed `--expert-cache 8000`, so the table is the only variable: | prompt | KL, IQ4_NL vs BF16 | top token | |---|---|---| | 2K | **0.156** | same | | 16K | **0.0030** | same | | 37K | **0.220** | same | Against your FP8 figures (0.019 at 2K, 0.0011 at 32K) that is **8-200x larger**, and above your reference point for the whole experts pack against an all-Q8_0 model (KL 0.0069 at 32K) by 20-30x. Two honest limits: the greedy **top token is unchanged in every case** - the deviation shows over sequences, not at one position - and these are single-position logits, not the answer-level KL you measured with a reference model's teacher-forced answers. ## The reader BF16 rows are 320 B rather than 90, and widening is exact: a bfloat16 is the top half of a float32, so `bf16_dequant_row` is a shift and a copy, little-endian bytes assembled explicitly so it does not depend on host endianness. No scale is read and no metadata is trusted - the values carry themselves. The direct reader already takes the table's own row size, so BF16 reads unbuffered like IQ4_NL and FP8 rather than falling back to the mapped path Q5_0 needs. (The comment claiming fixed 90-byte rows was stale from before FP8.) Tested in `ple_reader_test`: widening is exact over the patterns a lossy path breaks on - subnormals, +/-inf, both NaN encodings - and a synthetic BF16 GGUF reads back identically through Direct and Mmap. Both were mutation-checked: widening by `>>` instead of `<<`, and a 160 B row width, each fail the test. ## Cost None measurable. Same sweep, same card, same config, 256 output tokens per cell, cold: | prompt | IQ4_NL (90 B row) | BF16 (320 B row) | delta | |---:|---|---|---:| | 9,320 | 5,643 ms . 1,652 tok/s | 5,741 ms . 1,624 tok/s | -1.7% | | 15,558 | 7,003 . 2,222 | 7,167 . 2,171 | -2.3% | | 30,954 | 13,554 . 2,284 | 13,632 . 2,271 | -0.5% | | 57,731 | 25,131 . 2,297 | 25,185 . 2,292 | -0.2% | | 124,232 | 55,989 . 2,219 | 55,421 . 2,242 | +1.0% | | 232-237K | 121,713 . 1,945 | 119,884 . 1,937 | -0.4% | Within +/-2.3% across a 25x range, sign flipping - noise. 16 rows/token is dwarfed by streaming 47 GB of experts over PCIe. That matches your +1.8% at 32K on the 5090. ## Note on 0.1.38 Rebased onto 0.1.38, which moved this code twice: `max_inflight` 64 -> 256, and the Q5_0 `--ple-io direct` refusal removed with `row_bytes` generalised (#296). Both are kept as upstream has them; the BF16 checks now run inside upstream's parameterised `selftest(dir, rb)`. Mutation-checked after the rebase: `<< 16` -> `>> 8` fails the selftest on the first subnormal, so the BF16 coverage is live rather than merely present. ## Which to use - **IQ4_NL** - what ships today. ~8% off. - **F8_E4M3** - 51 GB, ~2.7% off. Below answer-level noise on your measurement, and the best cost/benefit for most installs. Build it from the FP8 checkpoint with this PR's packer. - **BF16** - 102 GB, exact. Worth it only when you need the values to be the checkpoint's exactly. Disk is the real cost: the table is streamed (`--ple-io direct`, never resident; the row cache is ~90 MB at the default), so this is disk and not RAM, but 51 GB or 102 GB has to exist.
No site
Links install, modelos, releases.