Pull requests / #586

#586 ple: read the PLE table at full BF16 precision (opt-in, alongside FP8)

closed · @constantindjonkam · 0 评论 · 在 GitHub 查看

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsWindows

描述

## What

Add a **BF16** PLE table as an **opt-in alternative** to the FP8 one. This is an addition, not a
replacement: FP8 stays the default path, nothing about its behaviour changes, and the existing
`tools/ple_fp8_pack.py` entry point keeps working.

A bfloat16 is the top half of a float32, so widening one is a shift — exact for every value, normal or
not. The checkpoint ships the table in F8_E4M3 (2.66% off per row) and in BF16 (the source of record), so
this reads the source instead of an approximation of it. Disk is the price: **102 GB against 51 GB**.

### What it does not do

- It does not change which table is used. `--ple-gguf` is unchanged; a table that is not a BF16 one takes
  exactly the path it took before. The only edit outside the reader and the packer is a comment.
- It does not replace `ple_fp8_pack.py`. That file is kept as a shim over the renamed
  `ple_table_pack.py`, because it is the only way to get a non-IQ4_NL table and command lines already
  exist against it.

## The measurement the previous review asked for

> "A teacher-forced comparison (KL and top-1 against the FP8 table) that shows a gain would be the reason
> to take it."

I built the FP8 table from `Qwen/Qwen3.8-Flash-Next-FP8` with the tool in `main` — 33 shards, 52.3 GB of
safetensors, packed to 51.2 GB `f8_e4m3` — and compared it against BF16 and against the IQ4_NL default.

**Teacher-forcing, decided from the dumps rather than from the text.** Window row 0 is
`P(next token | everything before pos)`, so two arms' rows at the same position are comparable only while
both have emitted the same tokens. That is decidable without a reference transcript: walk the windows in
order and stop at the first position where the greedy argmax differs. Every earlier row provably saw an
identical context.

An earlier version of this harness tried to align on the generated text and was wrong in an instructive
way: the engine writes timing lines to stdout too, so it compared log lines, called a *bit-identical*
control arm divergent, and compared post-divergence rows at equal positions — producing a KL of 19, which
reports that two contexts differ, not that two tables are 19 nats apart.

### Control first

8 prompts, 585 to 37K tokens, `--pcie-frac 0 --adapt-swaps 0` so residency is fixed and the table is the
only variable:

| arm | teacher-forced windows | KL median | KL p95 | KL max |
|---|---:|---:|---:|---:|
| **FP8 vs FP8 (control)** | **1701** | **0.00000** | **0.00000** | 0.00577 |
| BF16 vs FP8 | 30 | 0.01105 | 0.16647 | 0.50054 |
| IQ4_NL vs FP8 | 53 | 0.02280 | 0.36905 | 0.42025 |

The control is exactly zero at the median and the 95th percentile, with a single window out of 1701 above
it. So the BF16 row is a real difference — its median is about twice the control's *worst* window — and so
is IQ4_NL's, at roughly twice BF16's.

### The number that matters for a table choice

Verify windows of identical greedy output before the first disagreement, and how many of the 8 prompts
produced byte-identical output:

| pair | mean identical windows | identical full runs |
|---|---:|---:|
| **FP8 vs FP8 (control)** | **212.6** | **7 / 8** |
| BF16 vs FP8 | 3.8 | **0 / 8** |
| IQ4_NL vs FP8 | 6.6 | **0 / 8** |

**FP8 and BF16 are not interchangeable.** They disagree within about four verify windows — roughly 16
generated tokens — on every prompt tried. This is not a tail-logits curiosity; the table choice changes
the text.

And the practical result for a stock install: **IQ4_NL is about twice as far from FP8 as BF16 is**, in both
median and p95. Going IQ4_NL → FP8 recovers more than going FP8 → BF16 does.

### What this does not show, stated plainly

**Different is not better.** This measures how far apart the tables are, not which answers are better:
BF16 is the checkpoint's exact values, so it is "more correct" only by construction, and nothing here is a
quality benchmark. A perplexity number against a labelled set would answer the actual question and is not
included. On the previous review's evidence — BF16-vs-FP8 median KL 0.00087 against a 0.00080 noise floor,
for twice the disk — **FP8 remains the sensible default**, and I am not claiming otherwise.

Two limits on the numbers themselves: the teacher-forced sample is small (30 and 53 windows) precisely
*because* the arms diverge early, so the p95 rests on few points; and one control window of 1701 differs
(0.00577), so the floor is not a hard zero.

## Cost

None measurable. Same sweep, same card, 256 output tokens per cell, cold — IQ4_NL 90 B rows against BF16
320 B rows:

| prompt | IQ4_NL | BF16 | delta |
|---:|---|---|---:|
| 9,320 | 5,643 ms . 1,652 tok/s | 5,741 ms . 1,624 tok/s | -1.7% |
| 15,558 | 7,003 . 2,222 | 7,167 . 2,171 | -2.3% |
| 30,954 | 13,554 . 2,284 | 13,632 . 2,271 | -0.5% |
| 57,731 | 25,131 . 2,297 | 25,185 . 2,292 | -0.2% |
| 124,232 | 55,989 . 2,219 | 55,421 . 2,242 | +1.0% |
| 232-237K | 121,713 . 1,945 | 119,884 . 1,937 | -0.4% |

Within ±2.3% over a 25x range with the sign flipping. 16 rows/token is dwarfed by streaming 47 GB of
experts over PCIe.

## The reader

BF16 rows are 320 B rather than 160, and widening assembles the little-endian bytes explicitly rather than
reinterpreting memory, so it does not depend on host endianness. No scale is read: the values carry
themselves. The direct reader already takes the table's own row size, so BF16 reads unbuffered like FP8
instead of falling back to the mapped path.

Tested in `ple_reader_test --selftest` for the patterns a lossy path breaks on — subnormals, ±inf, both NaN
encodings — plus a synthetic BF16 GGUF round-tripping identically through Direct and Mmap. Both claims
were mutation-checked after the 0.1.38 rebase: widening by `>>` instead of `<<` fails on the first
subnormal, and a 160 B row width fails the round-trip.

## Which to use

- **F8_E4M3** — 51 GB, ~2.7% off. **The default, and the right choice for most installs.** Build it with
  `tools/ple_fp8_pack.py` (or `ple_table_pack.py`).
- **BF16** — 102 GB, exact. For anyone who needs the table to be the checkpoint's values rather than an
  approximation of them, and has the disk.
- **IQ4_NL** — 28.8 GB, ~8% off; what ships today. The measurement above says FP8 is the better step up if
  disk allows.

The table streams (`--ple-io direct`, never resident; the row cache is ~90 MB at the default), so this is
disk and not RAM — but 51 GB or 102 GB has to exist.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。