Pull requests / #896
#896 ple: accept Q5_1 and Q8_0 n-gram tables
closed · @anon761 · 0 commentaires · Sur GitHub
AMD / HIPNVIDIA / CUDAModels & quants
Description
## What `PleTable` reads `per_layer_token_embd.weight` as IQ4_NL, Q5_0 (#296, OrcaRouter) or the FP8 table. A plain quantized GGUF puts another type there, and the model refused to start: ``` strata generate: per_layer_token_embd.weight is Q5_1, not IQ4_NL, Q5_0 or FP8 (I8) ``` This adds the two the ordinary GGUFs ship: - **Q5_1** (120 B/row): the Uncensored finetune's Q5_K_M table. The 5-bit codes are *unsigned* with a per-block fp16 minimum added (`dequantize_row_q5_1`), unlike Q5_0's fixed `-16` offset. - **Q8_0** (170 B/row): unsloth's GGUFs and Swift-1.5. `PLE_ROW_BYTES_MAX` grows to the Q8_0 row; the row size, the type check, the decode branch and `format()`/`close()` follow. The direct reader already takes `row_bytes` at runtime (the comment about the stale refusal is upstream's and stays), so no mmap fallback is needed. ## Verified A `Qwen3.8-Flash-Next-Uncensored-Q5_K_M` GGUF (Q5_1 table) loads and serves 262k tokens on 2× RTX 3090 with a layer split; before the change it exited at startup.
Sur le site
Liens install, modèles, releases.