Pull requests / #651

#651 ple: read Q8_0 and Q5_1 n-gram tables

closed · @anon761 · 0 comentarios · En GitHub

BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants

Descripción

## What

Several ordinary GGUFs of Qwen3.8-Flash-Next ship `per_layer_token_embd.weight` in a format the PLE reader refuses at start:

| file | PLE table |
| --- | --- |
| Unsloth `UD-Q6_K_XL` | Q8_0 (170-byte rows, ~54 GB) |
| Swift-1.5 `Q4_K_L` | Q8_0 |
| a community `Q5_K_M` | Q5_1 (120-byte rows) |

```
strata generate: per_layer_token_embd.weight is Q5_1, not IQ4_NL, Q5_0 or FP8 (I8)
```

The rest of these files already has a path in the engine (Q5_K/Q6_K/Q5_1/Q8_0 experts, the K-quant projections, `--compat-bf16`). For the Q5_K_M file the table was the only thing keeping it from starting (end to end below); the Q8_0 tables are covered by the parity test on the real files.

- `dequantize_q5_1` in `strata/artifact/dequant.hpp` (ggml's `dequantize_row_q5_1`: unsigned 5-bit codes times `d` plus the block minimum `m`); Q8_0 rows use the existing `dequantize_q8_0`
- `PleTable` decodes Q5_1 / Q8_0 rows on both readers. The direct reader already takes the row size at runtime (since the FP8 table), so no reader change was needed; `PLE_ROW_BYTES_MAX` grows to 170
- `--ple-io ram`: before `mlock`, 16 threads fault the table in. `mlock` alone brings a cold file in one page at a time from one thread; on a ZFS pool that read the 54 GB Q8_0 table at ~0.47 GiB/s (about two minutes), 16 threads at ~2 GiB/s (measured from a cold cache before the rebase onto current `main`; the container this was re-tested in cannot drop the file cache). The fallback when `mlock` fails keeps the pages faulted in instead of touching them a second time.

## Tests

- `dequant_q5_1_test` (new, CTest): 4096 synthetic blocks incl. the edge codes 0 and 31 and a negative minimum against ggml's `to_float`, bit-exact. No model, no GPU.
- `dequant_bf16_test` knows Q5_1.
- `ple_q5_parity` now takes a Q5_0, Q5_1 or Q8_0 table and checks **both** the mapped and the direct reader (single rows and a 16-row batch) against ggml's `to_float` on the real file:

```
Q5_1 PLE, mapped reader: 320001536 rows, max_abs 0.000e+00 PASS     (Q5_K_M, shard 1)
Q5_1 PLE, direct reader: 320001536 rows, max_abs 0.000e+00 PASS
Q8_0 PLE, mapped reader: 320001536 rows, max_abs 0.000e+00 PASS     (Swift-1.5 Q4_K_L, shard 2)
Q8_0 PLE, direct reader: 320001536 rows, max_abs 0.000e+00 PASS
Q8_0 PLE, mapped reader: 320001536 rows, max_abs 0.000e+00 PASS     (UD-Q6_K_XL, shard 3)
Q8_0 PLE, direct reader: 320001536 rows, max_abs 0.000e+00 PASS
```

## End to end

The Q5_K_M file (Q5_1 table) on 2x RTX 3090, layer split auto, 262K context, `--kv int8 --pcie-frac 0`; `main` refuses it at start:

| `--ple-io` | prompt 2K | prompt 8K | decode |
| --- | ---: | ---: | ---: |
| `ram` (table locked in 3.2 s, warm file cache) | 705 tok/s | 2116 tok/s | 85-87 tok/s |
| `direct` (default) | 703 tok/s | 2117 tok/s | 65-74 tok/s |

En el sitio

Enlaces a install, modelos, releases.