Pull requests / #291

#291 PLE: the n-gram table in FP8 E4M3, as Qwen ships it (opt-in)

closed · @sergqwer · 0 comentarios · En GitHub

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsWindows

Descripción

## Summary

The engine reads the 51.2e9-value n-gram table only as IQ4_NL (ISTA-DASLab's shard 2, 90 B a row). That table is
8.1% off the checkpoint's own values per row (correlation 0.996-0.997 on rows from all 128 shards, so it is the same
table in the same order). Qwen ships it in FP8 E4M3: 128 shards of [2500012, 160] plus one scale. NVIDIA's NVFP4
checkpoint keeps the same FP8 table.

- `tools/ple_fp8_pack.py` copies those FP8 bytes, unchanged, into a one-tensor GGUF: `per_layer_token_embd.weight`
  [160, 320001536] as I8, with `strata.ple.format = f8_e4m3` and `strata.ple.scale`. The output is 51.2 GB, streamed
  at flat memory.
- `PleTable` opens either format; the row size (90 / 160 B) and the decode follow the file. `PleReader` takes the row
  size at open, and its cache sizes itself by it.
- A PLE-only file (`general.architecture = strata-ple`) is not added to the default dense-projection sources.
- The start line names the table's format.

## Measured

Measured on current main (0.1.29, d6708a4) with ISTA-DASLab's IQ2_XS, RTX 5090 32 GB, Ryzen 9 9950X3D, 128 GB DDR5, Windows 11. KL is first-token KL divergence (`STRATA_DUMP_FIRST_LOGITS`, #276); two runs of main itself differ by 0 on 5 of 6 short prompts and by 0.00013 on the sixth.

First-token KL of the IQ4_NL table against the FP8 one: 0.0009 (1K), 0.0013 (2K), 0.0002 (4K), 0.022 (32K); mean 0.0061, top-1 4/4.

- **Speed.** IQ2_XS decode with the FP8 table: 141 / 144 / 151 tok/s, against 136 / 143 / 140 with IQ4_NL
  (400-token answers, `--vram-reserve-mib 1500`). On an NVFP4 pack: 108-119 against 115-118.
- **Correctness.** `ple_fp8_parity`: rows through Direct, Mmap and `gather_batch` match torch's decode of the
  checkpoint bit for bit.

## Switch

Opt-in: `--ple-gguf <the FP8 GGUF>`. Nothing changes otherwise.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

En el sitio

Enlaces a install, modelos, releases.