Pull requests / #722

#722 Add Hadamard-INT2 GGUF support

closed · @SaibaWaipu · 0 comentários · No GitHub

NVIDIA / CUDAModels & quants

Descrição

## Summary
- Add a Hadamard-INT2 GGUF converter, reference codec, and format documentation, including metadata and expert-set validation.
- Add native packed-expert support with CPU and CUDA decode paths, plus Hadamard handling in prefill.
- Update GGUF tooling and document the runtime workflow.

## Validation
- Built the CPU target and `strata-gguf`.
- Ran synthetic single-shard and three-shard GGUF checks covering F16/BF16/F32 inputs, metadata and alignment, preservation of non-target payloads, importance matrices, transform dot invariance, and invalid inputs.
- CUDA was not compiled or runtime-tested because a CUDA toolchain was unavailable. Real-model quality and performance were not evaluated.

No site

Links install, modelos, releases.