Issues / #1611

#1611 [Bug]: a tensor-less GGUF shard whose aligned data start is past EOF is refused (unsloth UD-Q5_K_XL shard 1)

open · @CYoung83 · 0 comments · View on GitHub

NVIDIA / CUDAModels & quants

Description

### What happened

```
--native .../UD-Q5_K_XL/Qwen3.8-Flash-Next-UD-Q5_K_XL-00001-of-00006.gguf
strata generate: --native ...UD-Q5_K_XL-00001-of-00006.gguf: GGUF: data section starts past EOF
```

Shard 1 of unsloth/Qwen3.8-Flash-Next-GGUF `UD-Q5_K_XL` holds metadata only (0 tensors) and is 10,946,618 bytes.
Its 32-byte-aligned data start is 10,946,624, which is 6 bytes past EOF. Shard 1 of `UD-Q4_K_XL` and of `Q8_0` is
10,946,624 bytes, so its data start is exactly EOF, and it loads.

`include/strata/artifact/gguf_reader.hpp:458` throws whenever the aligned data start is past EOF, whether or not the
file has tensors. That line is unchanged on main `fb58e0d` (v0.1.41). llama.cpp's gguf-py at the pinned `3cf0325`
reads the file: `GGUFReader` reports 0 tensors and `data_offset` 10,946,624.

Workaround: pad a copy of shard 1 with 6 zero bytes (`truncate -s 10946624`) and symlink shards 2-6 beside it. The
engine then loads and runs the model.

A fix with a test follows in a PR.

### Strata version, GPU, OS

0.1.40.3 (source build, -DSTRATA_Q6K_EXPERTS=ON) · RTX 5090 32GB · Ubuntu 26.04, CUDA 13.3

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.