Pull requests / #619

#619 Add support for CYBER-FROST-3.8 (Blackfrost): read qwen4exp.nextn_predict_layers

closed · @mw00 · 0 コメント · GitHub で見る

AMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows

本文

# Add support for CYBER-FROST-3.8 (Blackfrost) — a qwen4exp fine-tune with its own nextn/MTP block

**Model:** [`peasantsmith/CYBER-FROST-3.8-PS-GUFF`](https://huggingface.co/peasantsmith/CYBER-FROST-3.8-PS-GUFF)
(a Q5_K_M requant of `Blackfrost-AI/CYBER-FROST-3.8-BF16`)
**Base:** upstream `main`, verified against `99f3dbd`
**Diff:** 3 files, 41 insertions(+), 6 deletions(-)

---

## What this adds

Strata currently runs **one** model family, `Qwen3.8-Flash-Next`, and the MTP machinery is built
around that family's files: `tools/mtp_fetch.py` fetches the draft head from
`Qwen/Qwen3.8-Flash-Next` at a pinned revision, and nothing in the engine reads a model's own
declaration of whether it has an MTP block.

**`CYBER-FROST-3.8` is a fine-tune of that family which declares its own nextn block in its GGUF
metadata:**

```
qwen4exp.block_count            = 49      <- 48 trunk + 1 prediction block
qwen4exp.nextn_predict_layers   = 1
```

**It cannot be loaded by the current engine.** This patch makes the engine read that declaration,
which is what any `qwen4exp` fine-tune carrying its own nextn block needs — the same class of
contribution as the existing `Swift 1.5` and OrcaRouter entries in `docs/MODELS.md`, but one that
does need engine changes.

## Why the engine refuses it today — four paths, four failures

| path | failure |
|---|---|
| **architecture check** — `include/strata/artifact/gguf_reader.hpp` | `block_count` (49) includes the MTP block, so the comparison against `Qwen4ExpGuard::block_count` (48) fails and **the model is rejected outright** |
| **dense loading** — `src/core/native_dense.cpp` | the three 2-D nextn projections are not in the dense-eligible suffix list, so they land on the packed fallback as a shape-only row and the loader aborts |
| **packing** — `tools/iq_pack.py` | `n_layers = 1 + max(blk.N)` counts the nextn block and emits an expert row the engine rejects as `"a malformed line"` (its parser bounds-checks `l < n_layers`) |

## The change

**1. `include/strata/artifact/gguf_reader.hpp`** — read `qwen4exp.nextn_predict_layers`
(absent → 0). `Qwen4ExpGuard` gains a `nextn` field, and the architecture check subtracts it from
the observed `block_count` before comparing, because **the kernels' contract is the trunk**:
nextn blocks sit after it and are a separate subsystem. A file reporting `block_count < nextn`
returns a precise error instead of a confusing mismatch.

**2. `src/core/native_dense.cpp`** — add the three 2-D nextn projections to the dense-eligible
suffix list:

```
.nextn.eh_proj.weight
.nextn.hc_head_down.weight
.nextn.hc_head_up.weight
```

The 1-D nextn norms (`enorm` / `hnorm` / `hc_head_norm`) are deliberately **not** listed — the
existing `shape.size() == 2` filter excludes them, so they stay on the packed fallback. **The
list is additive and every new entry passes the same guards the existing entries do** (`blk.`
prefix, supported quant type, 2-D shape), so **eligibility for any other model is unchanged.**

**3. `tools/iq_pack.py`** — subtract the declared `nextn` from `n_layers` so the pack's
`native_experts.txt` covers the trunk exactly. The drafter loads its own experts from the runtime
directory, so emitting a nextn row only makes the pack unloadable.

## Relationship to the existing MTP flow

**This does not replace or change `tools/mtp_fetch.py`, `mtp_pack.py` or `mtp_rt.py`.** Those stay
exactly as they are, and remain the way a draft head is obtained.

The gap is narrower: `mtp_fetch.py` is pinned to `Qwen/Qwen3.8-Flash-Next`, so a fine-tune whose
GGUF declares its own `nextn_predict_layers` can be *packed* with a head from that base — but the
engine then refuses to *load* it, because nothing reads the declaration. **These three changes
close that gap and nothing else.**

## Verification

**Applies and builds against current upstream:**

```
base                       : 99f3dbd (upstream main)
git apply --check          : clean
cmake -DSTRATA_ENABLE_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=70 ...
    configure              : RC=0
    build                  : RC=0, 0 errors
    binary                 : 42,428,592 bytes
```

**End-to-end load of the model this PR is for:**

```
strata generate: 1497 MiB of weights loaded from the pack
strata generate: 310 native projection matrices, 2314.01 MiB of weights
strata mtp: draft layer loaded, 799 MiB of VRAM (experts 675, dense 111)
strata generate: session is up (engine 0.1.38)
```

**The architecture guard passes, the MTP head binds, and the nextn projections load natively.**
Before the patch this model is rejected at the architecture check.

**The pack is trunk-correct:**

```
native_experts.txt     : 48 rows (layers 0..47)
references to blk.48   : 0
source declares        : block_count 49, nextn_predict_layers 1
```

**Speculative decoding works: the MTP drafter accepted 9,258 of 10,809 drafts (85.7%)** across the
measurement runs — 42% of the 274 drafting windows accepted every draft. Two drafters exist and
they are separate: the **MTP/suffix drafter** is the one bound by this patch, and its rate is the
85.7% above; the `--spec` verify window is a different mechanism and reports its own rate.

**No hardcoded paths, GPU counts, or host assumptions in the diff.**

## Draft-head provenance — stated plainly

The draft head used in these measurements was packed from the **base** model, because this
fine-tune ships no MTP weights of its own — **the same reason `mtp_fetch.py` reads the base
checkpoint.** It works — the MTP drafter accepts 85.7% of drafts — and speculative decoding is
**output-equivalent by
construction**, so generated text is unaffected.

**It is not necessarily the ideal head for this fine-tune** — one trained against the fine-tune's
distribution would accept more tokens. That is a property of the available weights, not of these
changes.

関連リンク

インストール・モデル・リリースへの站内リンク。