Pull requests / #619
#619 Add support for CYBER-FROST-3.8 (Blackfrost): read qwen4exp.nextn_predict_layers
closed · @mw00 · 0 评论 · 在 GitHub 查看
AMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows
描述
# Add support for CYBER-FROST-3.8 (Blackfrost) — a qwen4exp fine-tune with its own nextn/MTP block
**Model:** [`peasantsmith/CYBER-FROST-3.8-PS-GUFF`](https://huggingface.co/peasantsmith/CYBER-FROST-3.8-PS-GUFF)
(a Q5_K_M requant of `Blackfrost-AI/CYBER-FROST-3.8-BF16`)
**Base:** upstream `main`, verified against `99f3dbd`
**Diff:** 3 files, 41 insertions(+), 6 deletions(-)
---
## What this adds
Strata currently runs **one** model family, `Qwen3.8-Flash-Next`, and the MTP machinery is built
around that family's files: `tools/mtp_fetch.py` fetches the draft head from
`Qwen/Qwen3.8-Flash-Next` at a pinned revision, and nothing in the engine reads a model's own
declaration of whether it has an MTP block.
**`CYBER-FROST-3.8` is a fine-tune of that family which declares its own nextn block in its GGUF
metadata:**
```
qwen4exp.block_count = 49 <- 48 trunk + 1 prediction block
qwen4exp.nextn_predict_layers = 1
```
**It cannot be loaded by the current engine.** This patch makes the engine read that declaration,
which is what any `qwen4exp` fine-tune carrying its own nextn block needs — the same class of
contribution as the existing `Swift 1.5` and OrcaRouter entries in `docs/MODELS.md`, but one that
does need engine changes.
## Why the engine refuses it today — four paths, four failures
| path | failure |
|---|---|
| **architecture check** — `include/strata/artifact/gguf_reader.hpp` | `block_count` (49) includes the MTP block, so the comparison against `Qwen4ExpGuard::block_count` (48) fails and **the model is rejected outright** |
| **dense loading** — `src/core/native_dense.cpp` | the three 2-D nextn projections are not in the dense-eligible suffix list, so they land on the packed fallback as a shape-only row and the loader aborts |
| **packing** — `tools/iq_pack.py` | `n_layers = 1 + max(blk.N)` counts the nextn block and emits an expert row the engine rejects as `"a malformed line"` (its parser bounds-checks `l < n_layers`) |
## The change
**1. `include/strata/artifact/gguf_reader.hpp`** — read `qwen4exp.nextn_predict_layers`
(absent → 0). `Qwen4ExpGuard` gains a `nextn` field, and the architecture check subtracts it from
the observed `block_count` before comparing, because **the kernels' contract is the trunk**:
nextn blocks sit after it and are a separate subsystem. A file reporting `block_count < nextn`
returns a precise error instead of a confusing mismatch.
**2. `src/core/native_dense.cpp`** — add the three 2-D nextn projections to the dense-eligible
suffix list:
```
.nextn.eh_proj.weight
.nextn.hc_head_down.weight
.nextn.hc_head_up.weight
```
The 1-D nextn norms (`enorm` / `hnorm` / `hc_head_norm`) are deliberately **not** listed — the
existing `shape.size() == 2` filter excludes them, so they stay on the packed fallback. **The
list is additive and every new entry passes the same guards the existing entries do** (`blk.`
prefix, supported quant type, 2-D shape), so **eligibility for any other model is unchanged.**
**3. `tools/iq_pack.py`** — subtract the declared `nextn` from `n_layers` so the pack's
`native_experts.txt` covers the trunk exactly. The drafter loads its own experts from the runtime
directory, so emitting a nextn row only makes the pack unloadable.
## Relationship to the existing MTP flow
**This does not replace or change `tools/mtp_fetch.py`, `mtp_pack.py` or `mtp_rt.py`.** Those stay
exactly as they are, and remain the way a draft head is obtained.
The gap is narrower: `mtp_fetch.py` is pinned to `Qwen/Qwen3.8-Flash-Next`, so a fine-tune whose
GGUF declares its own `nextn_predict_layers` can be *packed* with a head from that base — but the
engine then refuses to *load* it, because nothing reads the declaration. **These three changes
close that gap and nothing else.**
## Verification
**Applies and builds against current upstream:**
```
base : 99f3dbd (upstream main)
git apply --check : clean
cmake -DSTRATA_ENABLE_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=70 ...
configure : RC=0
build : RC=0, 0 errors
binary : 42,428,592 bytes
```
**End-to-end load of the model this PR is for:**
```
strata generate: 1497 MiB of weights loaded from the pack
strata generate: 310 native projection matrices, 2314.01 MiB of weights
strata mtp: draft layer loaded, 799 MiB of VRAM (experts 675, dense 111)
strata generate: session is up (engine 0.1.38)
```
**The architecture guard passes, the MTP head binds, and the nextn projections load natively.**
Before the patch this model is rejected at the architecture check.
**The pack is trunk-correct:**
```
native_experts.txt : 48 rows (layers 0..47)
references to blk.48 : 0
source declares : block_count 49, nextn_predict_layers 1
```
**Speculative decoding works: the MTP drafter accepted 9,258 of 10,809 drafts (85.7%)** across the
measurement runs — 42% of the 274 drafting windows accepted every draft. Two drafters exist and
they are separate: the **MTP/suffix drafter** is the one bound by this patch, and its rate is the
85.7% above; the `--spec` verify window is a different mechanism and reports its own rate.
**No hardcoded paths, GPU counts, or host assumptions in the diff.**
## Draft-head provenance — stated plainly
The draft head used in these measurements was packed from the **base** model, because this
fine-tune ships no MTP weights of its own — **the same reason `mtp_fetch.py` reads the base
checkpoint.** It works — the MTP drafter accepts 85.7% of drafts — and speculative decoding is
**output-equivalent by
construction**, so generated text is unaffected.
**It is not necessarily the ideal head for this fine-tune** — one trained against the fine-tune's
distribution would accept more tokens. That is a property of the available weights, not of these
changes.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。