Pull requests / #140
#140 Add OrcaRouter Qwen3.8 Flash-Next Q4_K_S support
closed · @Suoriks · 0 コメント · GitHub で見る
Setup & installServer & APINVIDIA / CUDAModels & quantsDocumentationWindows
本文
## Changes Add support for the three original OrcaRouter Qwen3.8 Flash-Next Uncensored Q4_K_S GGUF shards. The loader finds experts whose tensors span shard boundaries. The packer keeps the Q4_K expert weights in GGUF form and reads the Q5_0 PLE table through mmap. CUDA and CPU paths handle the model's Q5_0/Q5_1 down projections and quantized small projections. `docs/ORCA_Q4_K_S.md` gives the setup commands and model-specific caveats. ## Checks - Built `strata`, `native_expert_parity`, `ple_q5_parity`, and `q5_projection_parity` from the rebased 0.1.24 tree on Windows, CUDA 12.6, L40S (sm_89). - `test_iq_pack.py`: 9 tests, 1 skipped. - `serve.test_server`: 38 tests passed. - PLE parity: 320,001,536 Q5_0 rows, zero maximum absolute difference. - Projection parity: 2,560 Q5_1 elements, zero maximum absolute difference and zero BF16 differences. - Native expert parity: layers 0, 5, 6, 20, and 47 passed on CUDA with the three source shards. These cover both Q5_0 and Q5_1 down projections and shard boundaries. The 0.1.23-based version also passed full text, vision, MTP, and 220K-token marker retrieval on two L40S cards. The 0.1.24 rebase has been rebuilt and checked with the parity tests above.
関連リンク
インストール・モデル・リリースへの站内リンク。