Pull requests / #91

#91 GGUF reader: refuse a duplicate tensor name at open, naming the tensor and the file

closed · merged 2026-09-29 · @Avicennasis · 0 Kommentare · Auf GitHub

Setup & installAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

Beschreibung

One fail-closed change on the GGUF reader, with a test that runs without a model or a GPU. (Companion PR: setup names a missing or short shard and `verify` reads back the manifest sha256.)

## `gguf_reader.hpp`: a duplicate tensor name is refused at open

`GgufFile::find()` is first-match over the tensor directory. A GGUF that carries one name twice therefore resolves that name to whichever entry came first; the other tensor's bytes are never read and nothing says so. GGUF has no index that could arbitrate between the two, so refusal is the whole fix: `parse()` keeps a set of the names it has seen and throws on the second sighting, naming both the tensor and the file:

```
GGUF: duplicate tensor name 'blk.0.attn_q.weight' in <path>
```

(llama.cpp's `gguf_init_from_file` refuses the same way.) `find()` itself is unchanged - the directory it walks is now unique.

Reproduction: `gguf_reader_test` (new, registered under `STRATA_BUILD_TESTS` beside `strata-gguf`, links only `strata_artifact`) writes a minimal GGUF v3 by hand - no metadata, two F32[8] tensors, 32-byte alignment - and opens it with the reader:

- distinct names: opens, `tensors().size() == 2`, `find()` sees both;
- the same name twice: refused at open; the error names the tensor and the file.

On `main` the second file opens and 3 of the 4 checks fail; with the change all 4 pass. The fixture was cross-checked against gguf-py's `GGUFReader` at the pinned llama.cpp commit and against `strata-gguf` (`tensors 2, data_start 128, size 192`). The file is closed before it is removed, for Windows.

## What I ran, and where

Linux x86_64 (Ubuntu 24.04, AMD Ryzen 7 5800X - no AVX-512, no NVIDIA GPU), gcc 13.3.0, CMake 3.28.3, Python 3.12.3. The CUDA targets were not built (`STRATA_ENABLE_CUDA=OFF`); this is a claim about my machine, not yours.

- `cmake -B build-cpu -DSTRATA_ENABLE_CUDA=OFF -DSTRATA_BUILD_TESTS=ON -DSTRATA_GGML_DIR=<llama.cpp at 3cf0325>` - the published tree has no `tests/`, so I added an empty `tests/CMakeLists.txt` locally (not committed), and `dequant_bf16_test` links `CUDA::cudart` outside the CUDA guard, so a stub imported target was injected for configure only. Both are pre-existing; not touched here.
- `cmake --build build-cpu -j8 -- -k`: every target built except `dequant_bf16_test` (`cuda_runtime.h` absent - expected without a toolkit; it is not a ctest).
- `ctest`: 9 tests; 8 passed. `gguf_reader_test` passed (4/4 rows). `expert_multi_test` fails on this CPU with the kernel's own refusal (`missing AVX512F ... AVX512-VBMI`); the branch touches nothing under `src/kernels`.
- `git diff --check` clean.

Branch is on `c1e9033` (0.1.20); happy to rebase if it has moved by the time you look.

Mehr auf der Site

Links zu Install, Modellen, Releases.