Pull requests / #67
#67 Support OrcaRouter IQ3_XXS with explicit BF16 compatibility packing
closed · merged 2026-09-28 · @acrogenesis · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installNVIDIA / CUDAModels & quantsLinux
描述
OrcaRouter's Qwen3.8 Flash Next Uncensored IQ3_XXS has the supported Qwen4Exp geometry, but quantizes small projections that Strata expects as BF16. This adds an explicit `tools/iq_pack.py --compat-bf16` workflow so that checkpoint can run with its own tokenizer and packed weights. - Dequantize only the required small projections and round to BF16 using nearest-even. Record the converted tensors in `compat-bf16.json`; leave expert weights, native attention weights, embeddings, and the disk-backed PLE table unchanged. - Preserve a converted BF16 PLE key in the native loader; keep the existing native Q2_0 key path. - Keep Hugging Face snapshot shard names when following symlinks, and reject missing shards, truncated tensors, and duplicate tensor names. - Add focused packing tests and manual setup documentation. The installer menu is unchanged; IQ3_XXS is the validated Orca quantization. Conversion adds BF16 rounding and does not restore the original checkpoint's precision. Validation on Linux with an RTX 5090, Ryzen 9 9950X3D, and 128 GB RAM: - Release CUDA build for sm_120; all 8 packing tests and 23 server tests pass. - GPU sampler, gated-residual, and KV-streaming parity checks pass. - Packed the complete two-shard model: 460 converted tensors, 1.39 GiB of converted BF16 weights. - Real inference completed all 12 throughput requests, including three with 261,880 input tokens and 256 output tokens in a 262,144-token window, without truncation or prompt reuse. [Benchmark write-up and raw measurements](https://acrogenesis.com/strata-vs-unsloth-qwen3-8-flash-next-rtx-5090/).
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。