Pull requests / #722
#722 Add Hadamard-INT2 GGUF support
closed · @SaibaWaipu · 0 评论 · 在 GitHub 查看
描述
## Summary - Add a Hadamard-INT2 GGUF converter, reference codec, and format documentation, including metadata and expert-set validation. - Add native packed-expert support with CPU and CUDA decode paths, plus Hadamard handling in prefill. - Update GGUF tooling and document the runtime workflow. ## Validation - Built the CPU target and `strata-gguf`. - Ran synthetic single-shard and three-shard GGUF checks covering F16/BF16/F32 inputs, metadata and alignment, preservation of non-target payloads, importance matrices, transform dot invariance, and invalid inputs. - CUDA was not compiled or runtime-tested because a CUDA toolchain was unavailable. Real-model quality and performance were not evaluated.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。