Pull requests / #586
#586 ple: read the PLE table at full BF16 precision (opt-in, alongside FP8)
closed · @constantindjonkam · 0 Kommentare · Auf GitHub
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsWindows
Beschreibung
## What Add a **BF16** PLE table as an **opt-in alternative** to the FP8 one. This is an addition, not a replacement: FP8 stays the default path, nothing about its behaviour changes, and the existing `tools/ple_fp8_pack.py` entry point keeps working. A bfloat16 is the top half of a float32, so widening one is a shift — exact for every value, normal or not. The checkpoint ships the table in F8_E4M3 (2.66% off per row) and in BF16 (the source of record), so this reads the source instead of an approximation of it. Disk is the price: **102 GB against 51 GB**. ### What it does not do - It does not change which table is used. `--ple-gguf` is unchanged; a table that is not a BF16 one takes exactly the path it took before. The only edit outside the reader and the packer is a comment. - It does not replace `ple_fp8_pack.py`. That file is kept as a shim over the renamed `ple_table_pack.py`, because it is the only way to get a non-IQ4_NL table and command lines already exist against it. ## The measurement the previous review asked for > "A teacher-forced comparison (KL and top-1 against the FP8 table) that shows a gain would be the reason > to take it." I built the FP8 table from `Qwen/Qwen3.8-Flash-Next-FP8` with the tool in `main` — 33 shards, 52.3 GB of safetensors, packed to 51.2 GB `f8_e4m3` — and compared it against BF16 and against the IQ4_NL default. **Teacher-forcing, decided from the dumps rather than from the text.** Window row 0 is `P(next token | everything before pos)`, so two arms' rows at the same position are comparable only while both have emitted the same tokens. That is decidable without a reference transcript: walk the windows in order and stop at the first position where the greedy argmax differs. Every earlier row provably saw an identical context. An earlier version of this harness tried to align on the generated text and was wrong in an instructive way: the engine writes timing lines to stdout too, so it compared log lines, called a *bit-identical* control arm divergent, and compared post-divergence rows at equal positions — producing a KL of 19, which reports that two contexts differ, not that two tables are 19 nats apart. ### Control first 8 prompts, 585 to 37K tokens, `--pcie-frac 0 --adapt-swaps 0` so residency is fixed and the table is the only variable: | arm | teacher-forced windows | KL median | KL p95 | KL max | |---|---:|---:|---:|---:| | **FP8 vs FP8 (control)** | **1701** | **0.00000** | **0.00000** | 0.00577 | | BF16 vs FP8 | 30 | 0.01105 | 0.16647 | 0.50054 | | IQ4_NL vs FP8 | 53 | 0.02280 | 0.36905 | 0.42025 | The control is exactly zero at the median and the 95th percentile, with a single window out of 1701 above it. So the BF16 row is a real difference — its median is about twice the control's *worst* window — and so is IQ4_NL's, at roughly twice BF16's. ### The number that matters for a table choice Verify windows of identical greedy output before the first disagreement, and how many of the 8 prompts produced byte-identical output: | pair | mean identical windows | identical full runs | |---|---:|---:| | **FP8 vs FP8 (control)** | **212.6** | **7 / 8** | | BF16 vs FP8 | 3.8 | **0 / 8** | | IQ4_NL vs FP8 | 6.6 | **0 / 8** | **FP8 and BF16 are not interchangeable.** They disagree within about four verify windows — roughly 16 generated tokens — on every prompt tried. This is not a tail-logits curiosity; the table choice changes the text. And the practical result for a stock install: **IQ4_NL is about twice as far from FP8 as BF16 is**, in both median and p95. Going IQ4_NL → FP8 recovers more than going FP8 → BF16 does. ### What this does not show, stated plainly **Different is not better.** This measures how far apart the tables are, not which answers are better: BF16 is the checkpoint's exact values, so it is "more correct" only by construction, and nothing here is a quality benchmark. A perplexity number against a labelled set would answer the actual question and is not included. On the previous review's evidence — BF16-vs-FP8 median KL 0.00087 against a 0.00080 noise floor, for twice the disk — **FP8 remains the sensible default**, and I am not claiming otherwise. Two limits on the numbers themselves: the teacher-forced sample is small (30 and 53 windows) precisely *because* the arms diverge early, so the p95 rests on few points; and one control window of 1701 differs (0.00577), so the floor is not a hard zero. ## Cost None measurable. Same sweep, same card, 256 output tokens per cell, cold — IQ4_NL 90 B rows against BF16 320 B rows: | prompt | IQ4_NL | BF16 | delta | |---:|---|---|---:| | 9,320 | 5,643 ms . 1,652 tok/s | 5,741 ms . 1,624 tok/s | -1.7% | | 15,558 | 7,003 . 2,222 | 7,167 . 2,171 | -2.3% | | 30,954 | 13,554 . 2,284 | 13,632 . 2,271 | -0.5% | | 57,731 | 25,131 . 2,297 | 25,185 . 2,292 | -0.2% | | 124,232 | 55,989 . 2,219 | 55,421 . 2,242 | +1.0% | | 232-237K | 121,713 . 1,945 | 119,884 . 1,937 | -0.4% | Within ±2.3% over a 25x range with the sign flipping. 16 rows/token is dwarfed by streaming 47 GB of experts over PCIe. ## The reader BF16 rows are 320 B rather than 160, and widening assembles the little-endian bytes explicitly rather than reinterpreting memory, so it does not depend on host endianness. No scale is read: the values carry themselves. The direct reader already takes the table's own row size, so BF16 reads unbuffered like FP8 instead of falling back to the mapped path. Tested in `ple_reader_test --selftest` for the patterns a lossy path breaks on — subnormals, ±inf, both NaN encodings — plus a synthetic BF16 GGUF round-tripping identically through Direct and Mmap. Both claims were mutation-checked after the 0.1.38 rebase: widening by `>>` instead of `<<` fails on the first subnormal, and a 160 B row width fails the round-trip. ## Which to use - **F8_E4M3** — 51 GB, ~2.7% off. **The default, and the right choice for most installs.** Build it with `tools/ple_fp8_pack.py` (or `ple_table_pack.py`). - **BF16** — 102 GB, exact. For anyone who needs the table to be the checkpoint's values rather than an approximation of them, and has the disk. - **IQ4_NL** — 28.8 GB, ~8% off; what ships today. The measurement above says FP8 is the better step up if disk allows. The table streams (`--ple-io direct`, never resident; the row cache is ~90 MB at the default), so this is disk and not RAM — but 51 GB or 102 GB has to exist.
Mehr auf der Site
Links zu Install, Modellen, Releases.