Pull requests / #464
#464 ple: read the n-gram table at full BF16 precision
closed · @constantindjonkam · 0 comments · View on GitHub
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quants
Description
## What Stop reading the PLE n-gram table from a lossy quantization of it. The table reaches the engine as **IQ4_NL** - ISTA-DASLab's shard 2 - which is about **8% off per row** against the checkpoint's own values. The checkpoint ships the table in **F8_E4M3** with a single `weight_scale`, at 2.66%, and in **BF16**, which is the source of record. This PR adds a packer (`tools/ple_table_pack.py`, renamed from `ple_fp8_pack.py`) that copies whichever form the checkpoint has, byte for byte - nothing is decoded, rounded or re-quantized - and a reader for the BF16 form. Thank you for the measurement on 0.1.37 / RTX 5090, and for building the BF16 table with this PR's tool against OrcaRouter's checkpoint. Two things from it that this PR now leans on: - **FP8 is already a good table.** Median KL over whole answers 0.00087 against BF16, with the same BF16 table at another cache size giving 0.00080 - i.e. below run-to-run noise - for 51 GB against 102 GB. That is the right call for your fork and the reason this PR leads with the packer rather than with BF16. - **Reading the shipped FP8 bytes is the faithful operation, and a cast would not be.** Your observation that only 6% of the checkpoint's BF16 values have the low 4 mantissa bits zero shows the FP8 table is a real re-quantization, not an upcast - which is exactly why the packer copies bytes instead of casting. ## What the review did not measure, and what we measured instead Your comparison is FP8 against BF16. The table a stock install actually runs is **IQ4_NL**, three times further out again, so that is the gap worth closing. First-window logits, same engine, same prompt, fixed `--expert-cache 8000`, so the table is the only variable: | prompt | KL, IQ4_NL vs BF16 | top token | |---|---|---| | 2K | **0.156** | same | | 16K | **0.0030** | same | | 37K | **0.220** | same | Against your FP8 figures (0.019 at 2K, 0.0011 at 32K) that is **8-200x larger**, and above your reference point for the whole experts pack against an all-Q8_0 model (KL 0.0069 at 32K) by 20-30x. Two honest limits: the greedy **top token is unchanged in every case** - the deviation shows over sequences, not at one position - and these are single-position logits, not the answer-level KL you measured with a reference model's teacher-forced answers. ## The reader BF16 rows are 320 B rather than 90, and widening is exact: a bfloat16 is the top half of a float32, so `bf16_dequant_row` is a shift and a copy, little-endian bytes assembled explicitly so it does not depend on host endianness. No scale is read and no metadata is trusted - the values carry themselves. The direct reader already takes the table's own row size, so BF16 reads unbuffered like IQ4_NL and FP8 rather than falling back to the mapped path Q5_0 needs. (The comment claiming fixed 90-byte rows was stale from before FP8.) Tested in `ple_reader_test`: widening is exact over the patterns a lossy path breaks on - subnormals, +/-inf, both NaN encodings - and a synthetic BF16 GGUF reads back identically through Direct and Mmap. Both were mutation-checked: widening by `>>` instead of `<<`, and a 160 B row width, each fail the test. ## Cost None measurable. Same sweep, same card, same config, 256 output tokens per cell, cold: | prompt | IQ4_NL (90 B row) | BF16 (320 B row) | delta | |---:|---|---|---:| | 9,320 | 5,643 ms . 1,652 tok/s | 5,741 ms . 1,624 tok/s | -1.7% | | 15,558 | 7,003 . 2,222 | 7,167 . 2,171 | -2.3% | | 30,954 | 13,554 . 2,284 | 13,632 . 2,271 | -0.5% | | 57,731 | 25,131 . 2,297 | 25,185 . 2,292 | -0.2% | | 124,232 | 55,989 . 2,219 | 55,421 . 2,242 | +1.0% | | 232-237K | 121,713 . 1,945 | 119,884 . 1,937 | -0.4% | Within +/-2.3% across a 25x range, sign flipping - noise. 16 rows/token is dwarfed by streaming 47 GB of experts over PCIe. That matches your +1.8% at 32K on the 5090. ## Note on 0.1.38 Rebased onto 0.1.38, which moved this code twice: `max_inflight` 64 -> 256, and the Q5_0 `--ple-io direct` refusal removed with `row_bytes` generalised (#296). Both are kept as upstream has them; the BF16 checks now run inside upstream's parameterised `selftest(dir, rb)`. Mutation-checked after the rebase: `<< 16` -> `>> 8` fails the selftest on the first subnormal, so the BF16 coverage is live rather than merely present. ## Which to use - **IQ4_NL** - what ships today. ~8% off. - **F8_E4M3** - 51 GB, ~2.7% off. Below answer-level noise on your measurement, and the best cost/benefit for most installs. Build it from the FP8 checkpoint with this PR's packer. - **BF16** - 102 GB, exact. Worth it only when you need the values to be the checkpoint's exactly. Disk is the real cost: the table is streamed (`--ple-io direct`, never resident; the row cache is ~90 MB at the default), so this is disk and not RAM, but 51 GB or 102 GB has to exist.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.