Pull requests / #67

#67 Support OrcaRouter IQ3_XXS with explicit BF16 compatibility packing

closed · merged 2026-09-28 · @acrogenesis · 0 コメント · GitHub で見る

BenchmarksSetup & installNVIDIA / CUDAModels & quantsLinux

本文

OrcaRouter's Qwen3.8 Flash Next Uncensored IQ3_XXS has the supported Qwen4Exp geometry, but quantizes small projections that Strata expects as BF16. This adds an explicit `tools/iq_pack.py --compat-bf16` workflow so that checkpoint can run with its own tokenizer and packed weights.

- Dequantize only the required small projections and round to BF16 using nearest-even. Record the converted tensors in `compat-bf16.json`; leave expert weights, native attention weights, embeddings, and the disk-backed PLE table unchanged.
- Preserve a converted BF16 PLE key in the native loader; keep the existing native Q2_0 key path.
- Keep Hugging Face snapshot shard names when following symlinks, and reject missing shards, truncated tensors, and duplicate tensor names.
- Add focused packing tests and manual setup documentation. The installer menu is unchanged; IQ3_XXS is the validated Orca quantization. Conversion adds BF16 rounding and does not restore the original checkpoint's precision.

Validation on Linux with an RTX 5090, Ryzen 9 9950X3D, and 128 GB RAM:

- Release CUDA build for sm_120; all 8 packing tests and 23 server tests pass.
- GPU sampler, gated-residual, and KV-streaming parity checks pass.
- Packed the complete two-shard model: 460 converted tensors, 1.39 GiB of converted BF16 weights.
- Real inference completed all 12 throughput requests, including three with 261,880 input tokens and 256 output tokens in a 262,144-token window, without truncation or prompt reuse.

[Benchmark write-up and raw measurements](https://acrogenesis.com/strata-vs-unsloth-qwen3-8-flash-next-rtx-5090/).

関連リンク

インストール・モデル・リリースへの站内リンク。