Pull requests / #1612
#1612 gguf_reader: open a tensor-less shard that ends before its aligned data start
open · @CYoung83 · 0 comentários · No GitHub
Setup & installAMD / HIPNVIDIA / CUDAModels & quants
Descrição
## Title
Issue: Resolves #1611
## Summary
A GGUF file with no tensors (a split model's metadata-only first shard) whose header ends before the
32-byte-aligned data start was refused with "data section starts past EOF". It now opens. A file with
tensors and the same truncation is still refused.
## What changed
- include/strata/artifact/gguf_reader.hpp: when the aligned data start is past EOF, throw only if the
file has tensors. Otherwise clamp data_start_ to the file size. That keeps `file_size() - data_start()`
at 0 in the payload checks that compute it (native_dense.cpp, native_head.cpp, expert_source.cpp,
in_bounds), where it would otherwise wrap. Nothing else in the reader changes.
- src/artifact/gguf_reader_test.cpp: two synthetic cases.
- No tensors, file ends at its header: opens, and data_start() == file_size().
- One tensor, same truncation: refused with the same message.
## Extra Notes
Tested:
- On main fb58e0d with GCC 13.3 (-std=c++20 -Wall -Wextra, no warnings): gguf_reader_test fails 2 checks
without the header change and passes with it. gguf_split_test passes either way. Clang 18.1 (-std=c++17)
passes.
- On main fb58e0d (v0.1.41), CUDA build (sm_120, Release, -DSTRATA_Q6K_EXPERTS=ON), on an RTX 5090: the engine
as released against the engine with these two commits, run back to back.
- The fixed engine loads the original unsloth UD-Q5_K_XL files. The released engine gets a copy of shard 1
padded with 6 zero bytes.
- The rows STRATA_LOGPOS_TOPK=256 wrote for three ~570-token requests are byte-identical between the two.
The same holds for Q8_0, whose shard 1 needs no padding: 3,424 rows, all identical.
- The CUDA parity tests pass: native_expert_parity (including q6_K_q8_0), native_grouped_parity and
ple_reader_selftest.
- The same comparison on v0.1.40.3: also identical.
- HIP and SYCL, compile only (no AMD or Intel GPU here), on the released tree and on the tree with these
commits:
- HIP, in rocm/dev-ubuntu-24.04:7.2.4-complete with setup.py's build_engine_hip() options
(-DSTRATA_PREFILL_MMQ=ON, CMAKE_HIP_ARCHITECTURES=gfx1100;gfx1201): both trees build `strata` with no
errors, and gguf_reader_test and gguf_split_test pass. The fixed tree's reader test includes the two new
cases.
- SYCL, in strata-sycl-dev (sycl/tools/Dockerfile) with sycl/tools/build.sh, JIT target: both trees end
with BUILD EXIT 0 and errors: 0.
gguf_reader.hpp says it is generated from src/artifact/gguf_reader.cpp by scripts/split_artifact.py. That
.cpp no longer holds the parser, and scripts/ isn't in the published tree, so I edited the header directly.
No site
Links install, modelos, releases.