Pull requests / #865
#865 ngram: add opt-in Q8_0 PLE table support
closed · @CC-David-CC · 0 Kommentare · Auf GitHub
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows
Beschreibung
The PLE reader currently rejects Q8_0 tables even when the model's expert weights are supported. Add 170-byte Q8_0 rows through the existing scalar dequantizer and mapped reader. This is the independent Q8 follow-up discussed in #689. It starts from upstream `6f32ec0` and does not include or depend on the rotation patch in #864. - **Off by default:** requires exactly `STRATA_EXPERIMENTAL_Q8_PLE=1`. - Supports `--ple-io mmap` and `--ple-io ram`; direct Q8 PLE I/O is explicitly rejected. - Checks shape/size arithmetic and file bounds; clears state on close/reopen. Preserves existing prefetch behavior and other PLE formats. - No serving, scheduling, expert-placement, installer preset or README changes. ### Validation Fresh standalone build at `18d3da487f82c8e7b5d8e91f6fdd8c75972182f1`, RTX PRO 6000 Blackwell 96GB / Ryzen 7950X / GCC 13.3 / CUDA 13.2: - Both CTest checks passed, covering flag values, signed/scaled rows, mmap/RAM, malformed files and lifecycle. - Real Q8 PLE: 1,059 probes plus batch/issue-collect checks bit-identical to ggml. - Fresh native run: **1,024 generated token IDs match the retained rotation-off baseline**, after the same 1,024-token input. Zero offered drafts and prompt reuse, FP16 KV. The model uses Q8 experts/PLE, compatibility BF16 small projections and a Q5_K head. The EOS sentinel forces the count. This is a correctness check, not a quality or speedup claim. Q8 PLE row data alone is about 50.66 GiB; the full-model configuration requires substantial memory. AMD and Windows GPU execution remain untested. [Enablement and tests](https://github.com/CC-David-CC/Strata-a5500/blob/feat/q8-ple-reader/docs/Q8_PLE.md) | [Fresh receipt](https://github.com/CC-David-CC/Strata-a5500/blob/feat/q8-ple-reader/bench/results/q8-ple-reader-validation.json) | [Raw IDs](https://github.com/CC-David-CC/Strata-a5500/blob/feat/q8-ple-reader/bench/results/q8-ple-reader-token-ids.json) The guide also links the older combined branch's token-identical 1K/16K/32K rotation comparisons, clearly separated from this standalone validation.
Mehr auf der Site
Links zu Install, Modellen, Releases.