Pull requests / #896

#896 ple: accept Q5_1 and Q8_0 n-gram tables

closed · @anon761 · 0 コメント · GitHub で見る

AMD / HIPNVIDIA / CUDAModels & quants

本文

## What

`PleTable` reads `per_layer_token_embd.weight` as IQ4_NL, Q5_0 (#296, OrcaRouter) or the FP8 table. A plain quantized GGUF puts another type there, and the model refused to start:

```
strata generate: per_layer_token_embd.weight is Q5_1, not IQ4_NL, Q5_0 or FP8 (I8)
```

This adds the two the ordinary GGUFs ship:

- **Q5_1** (120 B/row): the Uncensored finetune's Q5_K_M table. The 5-bit codes are *unsigned* with a per-block fp16 minimum added (`dequantize_row_q5_1`), unlike Q5_0's fixed `-16` offset.
- **Q8_0** (170 B/row): unsloth's GGUFs and Swift-1.5.

`PLE_ROW_BYTES_MAX` grows to the Q8_0 row; the row size, the type check, the decode branch and `format()`/`close()` follow.

The direct reader already takes `row_bytes` at runtime (the comment about the stale refusal is upstream's and stays), so no mmap fallback is needed.

## Verified

A `Qwen3.8-Flash-Next-Uncensored-Q5_K_M` GGUF (Q5_1 table) loads and serves 262k tokens on 2× RTX 3090 with a layer split; before the change it exited at startup.

関連リンク

インストール・モデル・リリースへの站内リンク。