Pull requests / #140
#140 Add OrcaRouter Qwen3.8 Flash-Next Q4_K_S support
closed · @Suoriks · 0 comments · View on GitHub
Setup & installServer & APINVIDIA / CUDAModels & quantsDocumentationWindows
Description
## Changes Add support for the three original OrcaRouter Qwen3.8 Flash-Next Uncensored Q4_K_S GGUF shards. The loader finds experts whose tensors span shard boundaries. The packer keeps the Q4_K expert weights in GGUF form and reads the Q5_0 PLE table through mmap. CUDA and CPU paths handle the model's Q5_0/Q5_1 down projections and quantized small projections. `docs/ORCA_Q4_K_S.md` gives the setup commands and model-specific caveats. ## Checks - Built `strata`, `native_expert_parity`, `ple_q5_parity`, and `q5_projection_parity` from the rebased 0.1.24 tree on Windows, CUDA 12.6, L40S (sm_89). - `test_iq_pack.py`: 9 tests, 1 skipped. - `serve.test_server`: 38 tests passed. - PLE parity: 320,001,536 Q5_0 rows, zero maximum absolute difference. - Projection parity: 2,560 Q5_1 elements, zero maximum absolute difference and zero BF16 differences. - Native expert parity: layers 0, 5, 6, 20, and 47 passed on CUDA with the three source shards. These cover both Q5_0 and Q5_1 down projections and shard boundaries. The 0.1.23-based version also passed full text, vision, MTP, and 220K-token marker retrieval on two L40S cards. The 0.1.24 rebase has been rebuilt and checked with the parity tests above.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.