Pull requests / #464

#464 ple: read the n-gram table at full BF16 precision

closed · @constantindjonkam · 0 comments · View on GitHub

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quants

Description

## What

Stop reading the PLE n-gram table from a lossy quantization of it.

The table reaches the engine as **IQ4_NL** - ISTA-DASLab's shard 2 - which is about **8% off per row**
against the checkpoint's own values. The checkpoint ships the table in **F8_E4M3** with a single
`weight_scale`, at 2.66%, and in **BF16**, which is the source of record.

This PR adds a packer (`tools/ple_table_pack.py`, renamed from `ple_fp8_pack.py`) that copies whichever
form the checkpoint has, byte for byte - nothing is decoded, rounded or re-quantized - and a reader for
the BF16 form.

Thank you for the measurement on 0.1.37 / RTX 5090, and for building the BF16 table with this PR's tool
against OrcaRouter's checkpoint. Two things from it that this PR now leans on:

- **FP8 is already a good table.** Median KL over whole answers 0.00087 against BF16, with the same BF16
  table at another cache size giving 0.00080 - i.e. below run-to-run noise - for 51 GB against 102 GB.
  That is the right call for your fork and the reason this PR leads with the packer rather than with
  BF16.
- **Reading the shipped FP8 bytes is the faithful operation, and a cast would not be.** Your observation
  that only 6% of the checkpoint's BF16 values have the low 4 mantissa bits zero shows the FP8 table is
  a real re-quantization, not an upcast - which is exactly why the packer copies bytes instead of casting.

## What the review did not measure, and what we measured instead

Your comparison is FP8 against BF16. The table a stock install actually runs is **IQ4_NL**, three times
further out again, so that is the gap worth closing. First-window logits, same engine, same prompt,
fixed `--expert-cache 8000`, so the table is the only variable:

| prompt | KL, IQ4_NL vs BF16 | top token |
|---|---|---|
| 2K | **0.156** | same |
| 16K | **0.0030** | same |
| 37K | **0.220** | same |

Against your FP8 figures (0.019 at 2K, 0.0011 at 32K) that is **8-200x larger**, and above your
reference point for the whole experts pack against an all-Q8_0 model (KL 0.0069 at 32K) by 20-30x.

Two honest limits: the greedy **top token is unchanged in every case** - the deviation shows over
sequences, not at one position - and these are single-position logits, not the answer-level KL you
measured with a reference model's teacher-forced answers.

## The reader

BF16 rows are 320 B rather than 90, and widening is exact: a bfloat16 is the top half of a float32, so
`bf16_dequant_row` is a shift and a copy, little-endian bytes assembled explicitly so it does not depend
on host endianness. No scale is read and no metadata is trusted - the values carry themselves.

The direct reader already takes the table's own row size, so BF16 reads unbuffered like IQ4_NL and FP8
rather than falling back to the mapped path Q5_0 needs. (The comment claiming fixed 90-byte rows was
stale from before FP8.)

Tested in `ple_reader_test`: widening is exact over the patterns a lossy path breaks on - subnormals,
+/-inf, both NaN encodings - and a synthetic BF16 GGUF reads back identically through Direct and Mmap.
Both were mutation-checked: widening by `>>` instead of `<<`, and a 160 B row width, each fail the test.

## Cost

None measurable. Same sweep, same card, same config, 256 output tokens per cell, cold:

| prompt | IQ4_NL (90 B row) | BF16 (320 B row) | delta |
|---:|---|---|---:|
| 9,320 | 5,643 ms . 1,652 tok/s | 5,741 ms . 1,624 tok/s | -1.7% |
| 15,558 | 7,003 . 2,222 | 7,167 . 2,171 | -2.3% |
| 30,954 | 13,554 . 2,284 | 13,632 . 2,271 | -0.5% |
| 57,731 | 25,131 . 2,297 | 25,185 . 2,292 | -0.2% |
| 124,232 | 55,989 . 2,219 | 55,421 . 2,242 | +1.0% |
| 232-237K | 121,713 . 1,945 | 119,884 . 1,937 | -0.4% |

Within +/-2.3% across a 25x range, sign flipping - noise. 16 rows/token is dwarfed by streaming 47 GB
of experts over PCIe. That matches your +1.8% at 32K on the 5090.

## Note on 0.1.38

Rebased onto 0.1.38, which moved this code twice: `max_inflight` 64 -> 256, and the Q5_0
`--ple-io direct` refusal removed with `row_bytes` generalised (#296). Both are kept as upstream has them;
the BF16 checks now run inside upstream's parameterised `selftest(dir, rb)`. Mutation-checked after the
rebase: `<< 16` -> `>> 8` fails the selftest on the first subnormal, so the BF16 coverage is live rather
than merely present.

## Which to use

- **IQ4_NL** - what ships today. ~8% off.
- **F8_E4M3** - 51 GB, ~2.7% off. Below answer-level noise on your measurement, and the best
  cost/benefit for most installs. Build it from the FP8 checkpoint with this PR's packer.
- **BF16** - 102 GB, exact. Worth it only when you need the values to be the checkpoint's exactly.

Disk is the real cost: the table is streamed (`--ple-io direct`, never resident; the row cache is ~90 MB
at the default), so this is disk and not RAM, but 51 GB or 102 GB has to exist.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.