Pull requests / #865

#865 ngram: add opt-in Q8_0 PLE table support

closed · @CC-David-CC · 0 commentaires · Sur GitHub

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows

Description

The PLE reader currently rejects Q8_0 tables even when the model's expert weights are supported. Add 170-byte Q8_0 rows through the existing scalar dequantizer and mapped reader.

This is the independent Q8 follow-up discussed in #689. It starts from upstream `6f32ec0` and does not include or depend on the rotation patch in #864.

- **Off by default:** requires exactly `STRATA_EXPERIMENTAL_Q8_PLE=1`.
- Supports `--ple-io mmap` and `--ple-io ram`; direct Q8 PLE I/O is explicitly rejected.
- Checks shape/size arithmetic and file bounds; clears state on close/reopen. Preserves existing prefetch behavior and other PLE formats.
- No serving, scheduling, expert-placement, installer preset or README changes.

### Validation

Fresh standalone build at `18d3da487f82c8e7b5d8e91f6fdd8c75972182f1`, RTX PRO 6000 Blackwell 96GB / Ryzen 7950X / GCC 13.3 / CUDA 13.2:

- Both CTest checks passed, covering flag values, signed/scaled rows, mmap/RAM, malformed files and lifecycle.
- Real Q8 PLE: 1,059 probes plus batch/issue-collect checks bit-identical to ggml.
- Fresh native run: **1,024 generated token IDs match the retained rotation-off baseline**, after the same 1,024-token input. Zero offered drafts and prompt reuse, FP16 KV.

The model uses Q8 experts/PLE, compatibility BF16 small projections and a Q5_K head. The EOS sentinel forces the count. This is a correctness check, not a quality or speedup claim. Q8 PLE row data alone is about 50.66 GiB; the full-model configuration requires substantial memory. AMD and Windows GPU execution remain untested.

[Enablement and tests](https://github.com/CC-David-CC/Strata-a5500/blob/feat/q8-ple-reader/docs/Q8_PLE.md) | [Fresh receipt](https://github.com/CC-David-CC/Strata-a5500/blob/feat/q8-ple-reader/bench/results/q8-ple-reader-validation.json) | [Raw IDs](https://github.com/CC-David-CC/Strata-a5500/blob/feat/q8-ple-reader/bench/results/q8-ple-reader-token-ids.json)

The guide also links the older combined branch's token-identical 1K/16K/32K rotation comparisons, clearly separated from this standalone validation.

Sur le site

Liens install, modèles, releases.