Pull requests / #464

#464 ple: read the n-gram table at full BF16 precision

closed · @constantindjonkam · 0 commentaires · Sur GitHub

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quants

Description

## What

Stop reading the PLE n-gram table from a lossy quantization of it.

The table reaches the engine as **IQ4_NL** - ISTA-DASLab's shard 2 - which is about **8% off per row**
against the checkpoint's own values. The checkpoint ships the table in **F8_E4M3** with a single
`weight_scale`, at 2.66%, and in **BF16**, which is the source of record.

This PR adds a packer (`tools/ple_table_pack.py`, renamed from `ple_fp8_pack.py`) that copies whichever
form the checkpoint has, byte for byte - nothing is decoded, rounded or re-quantized - and a reader for
the BF16 form.

Thank you for the measurement on 0.1.37 / RTX 5090, and for building the BF16 table with this PR's tool
against OrcaRouter's checkpoint. Two things from it that this PR now leans on:

- **FP8 is already a good table.** Median KL over whole answers 0.00087 against BF16, with the same BF16
  table at another cache size giving 0.00080 - i.e. below run-to-run noise - for 51 GB against 102 GB.
  That is the right call for your fork and the reason this PR leads with the packer rather than with
  BF16.
- **Reading the shipped FP8 bytes is the faithful operation, and a cast would not be.** Your observation
  that only 6% of the checkpoint's BF16 values have the low 4 mantissa bits zero shows the FP8 table is
  a real re-quantization, not an upcast - which is exactly why the packer copies bytes instead of casting.

## What the review did not measure, and what we measured instead

Your comparison is FP8 against BF16. The table a stock install actually runs is **IQ4_NL**, three times
further out again, so that is the gap worth closing. First-window logits, same engine, same prompt,
fixed `--expert-cache 8000`, so the table is the only variable:

| prompt | KL, IQ4_NL vs BF16 | top token |
|---|---|---|
| 2K | **0.156** | same |
| 16K | **0.0030** | same |
| 37K | **0.220** | same |

Against your FP8 figures (0.019 at 2K, 0.0011 at 32K) that is **8-200x larger**, and above your
reference point for the whole experts pack against an all-Q8_0 model (KL 0.0069 at 32K) by 20-30x.

Two honest limits: the greedy **top token is unchanged in every case** - the deviation shows over
sequences, not at one position - and these are single-position logits, not the answer-level KL you
measured with a reference model's teacher-forced answers.

## The reader

BF16 rows are 320 B rather than 90, and widening is exact: a bfloat16 is the top half of a float32, so
`bf16_dequant_row` is a shift and a copy, little-endian bytes assembled explicitly so it does not depend
on host endianness. No scale is read and no metadata is trusted - the values carry themselves.

The direct reader already takes the table's own row size, so BF16 reads unbuffered like IQ4_NL and FP8
rather than falling back to the mapped path Q5_0 needs. (The comment claiming fixed 90-byte rows was
stale from before FP8.)

Tested in `ple_reader_test`: widening is exact over the patterns a lossy path breaks on - subnormals,
+/-inf, both NaN encodings - and a synthetic BF16 GGUF reads back identically through Direct and Mmap.
Both were mutation-checked: widening by `>>` instead of `<<`, and a 160 B row width, each fail the test.

## Cost

None measurable. Same sweep, same card, same config, 256 output tokens per cell, cold:

| prompt | IQ4_NL (90 B row) | BF16 (320 B row) | delta |
|---:|---|---|---:|
| 9,320 | 5,643 ms . 1,652 tok/s | 5,741 ms . 1,624 tok/s | -1.7% |
| 15,558 | 7,003 . 2,222 | 7,167 . 2,171 | -2.3% |
| 30,954 | 13,554 . 2,284 | 13,632 . 2,271 | -0.5% |
| 57,731 | 25,131 . 2,297 | 25,185 . 2,292 | -0.2% |
| 124,232 | 55,989 . 2,219 | 55,421 . 2,242 | +1.0% |
| 232-237K | 121,713 . 1,945 | 119,884 . 1,937 | -0.4% |

Within +/-2.3% across a 25x range, sign flipping - noise. 16 rows/token is dwarfed by streaming 47 GB
of experts over PCIe. That matches your +1.8% at 32K on the 5090.

## Note on 0.1.38

Rebased onto 0.1.38, which moved this code twice: `max_inflight` 64 -> 256, and the Q5_0
`--ple-io direct` refusal removed with `row_bytes` generalised (#296). Both are kept as upstream has them;
the BF16 checks now run inside upstream's parameterised `selftest(dir, rb)`. Mutation-checked after the
rebase: `<< 16` -> `>> 8` fails the selftest on the first subnormal, so the BF16 coverage is live rather
than merely present.

## Which to use

- **IQ4_NL** - what ships today. ~8% off.
- **F8_E4M3** - 51 GB, ~2.7% off. Below answer-level noise on your measurement, and the best
  cost/benefit for most installs. Build it from the FP8 checkpoint with this PR's packer.
- **BF16** - 102 GB, exact. Worth it only when you need the values to be the checkpoint's exactly.

Disk is the real cost: the table is streamed (`--ple-io direct`, never resident; the row cache is ~90 MB
at the default), so this is disk and not RAM, but 51 GB or 102 GB has to exist.

Sur le site

Liens install, modèles, releases.