Pull requests / #296
#296 Add OrcaRouter Q4_K_S support (0.1.30)
closed · @Suoriks · 0 评论 · 在 GitHub 查看
Setup & installAMD / HIPNVIDIA / CUDAModels & quantsWindows
描述
This rebases the OrcaRouter Q4_K_S port from #140 onto Strata 0.1.30. The new `STRATA_ORCA_Q4KS_MMQ` option defaults to `OFF`. When enabled, it adds Q4_K, Q5_0 and Q5_1 GGML MMQ instances to the CUDA build and to the optional HIP prefill build. The existing Q8_0 draft-layer instance stays in both configurations. The port also reads expert tensors that cross GGUF shard boundaries, handles the Q5_0 PLE table and Q5_1 projections, and adds manual setup instructions. Checks on Windows with two NVIDIA L40S cards: - CUDA Release `strata` build passed with the option both `ON` and `OFF`. Both generated build graphs retain Q8_0; only `ON` adds the three new instances. - Focused pack tests: 8 passed, 1 Windows symlink test skipped. The existing BF16 pack byte-identity test passed. - Real GGUF parity: Q5_0 PLE, Q5_1 projection and native experts at layers 0, 5, 6, 20 and 47 passed. The HIP sources use the same CMake instance list. This machine has no ROCm toolchain or AMD GPU, so I could not compile or run the HIP build here. CMake stops at HIP compiler discovery (`Failed to find HIP root directory`).
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。