Pull requests / #911

#911 setup: --inspect, a GGUF's real quantization and whether Strata runs it, from its headers only

closed · @1872183316 · 0 comentários · No GitHub

Setup & installNVIDIA / CUDAModels & quants

Descrição

`--inspect SOURCE [VARIANT]` says what a GGUF really holds and whether Strata runs it, from its headers only (a few MB,
also over the network), before anyone downloads 70-110 GB. And the size menu shows each size's real routed-expert bits.

```
./setup.sh --inspect ms:unsloth/Qwen3.8-Flash-Next-GGUF UD-Q2_K_XL      # ModelScope (hf:owner/repo for Hugging Face)
./setup.sh --inspect /data/models/some-file-00001-of-00002.gguf          # a file or a folder; a URL works too
```

It prints every weight group's parameters, size, real bits per weight and storage types (shares of bytes), then one of:

- **known**: the routed experts, the tensor count and the bytes match a file setup installs: "this is setup's
  --family qwen --model IQ3_XXS" (under any file name, e.g. a renamed or re-uploaded copy);
- **same layout**: the routed experts are stored exactly as in a known file, the rest differs (a fine-tune with the
  same recipe): "may run as that size; not tested";
- **no**: another architecture, or routed experts in a format no file Strata runs uses (e.g. Unsloth's UD-IQ1_S: IQ1_S);
- **untested**: experts in known formats, but not a file setup installs; it names every tensor stored in a format no
  known file has for it.

The menu line becomes e.g. `1) Q2_0  2-bit, the fastest; download 66 GB, uses ~34 GB of RAM, experts 2.25 bits/weight`
(IQ2_XS 2.35, IQ3_XXS 2.84, IQ3_S 3.33, Coder IQ1_M 3.33 over its kept experts, UD-IQ4_XS 3.94, UD-Q4_K_XL 5.10).

## How

- `tools/quantscope.py`: the header reader (GGUF, safetensors; ModelScope / Hugging Face / URL / local, HTTP range
  requests, standard library only), from https://github.com/1872183316/quantscope (dual-licensed MIT OR Apache-2.0,
  so it can be MIT here), with ggml type 42 named Q2_0.
- `tools/strata_inspect.py`: the comparison, and `--fingerprints OUT` to rebuild the table.
- `data/gguf_fingerprints.json` (8.6 KB): the 10 files setup installs, read from ModelScope's copies on 2026-10-05:
  architecture, tensor count, bytes, routed-expert bytes per type, and, per tensor role, the types any of them use.
- `setup.py`: `--inspect` runs the tool and exits; `expert_bits()` for the menu (no network at install time).

## One thing to confirm

`SUPPORTED_GGUFS` says Unsloth's UD-IQ3_XXS and UD-Q2_K_XL cannot be used. From the headers, their routed experts use
only formats the known files use (UD-Q2_K_XL: IQ2_XS, IQ3_XXS, IQ4_NL); what differs is `output.weight` (Q4_K) and
`token_embd.weight` (Q5_K; UD-IQ3_XXS: Q6_K). So `--inspect` says "untested" and names those two, rather than "no".
If the engine refuses those head / embedding formats, a line in `strata_inspect.py` can turn that into "no"; I did not
try to run either file.

## Tests

`tools/test_strata_inspect.py` (5, offline, on the fingerprint table): every installed file is known, a changed size is
"same layout", another architecture and an unknown expert format are "no", and an untested file names exactly its new
tensors. All `tools/test_setup_*.py` pass (the menu line is printed by the golden tests).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

No site

Links install, modelos, releases.