Pull requests / #54
#54 Support pruned expert variants (GSQ-RCO-Coder: 256 of 512 experts)
closed · merged 2026-09-28 · @pjgmobile · 0 comentários · No GitHub
BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsSecurity
Descrição
# Support pruned expert variants (GSQ-RCO-Coder: 256 of 512 experts) ## Motivation ISTA-DASLab recently published `Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF` — the same architecture with **half the routed experts pruned away per layer** (512 → 256, `qwen4exp.expert_count = 256`, still 10 active). The engine currently specialises to the canonical geometry in a number of places, so a pruned model fails at a different spot for each one: the pack tool divides by a hardcoded 512, the pack loader validates against the compile-time default, the runtime graph is default-constructed, the prefill kernels and buffers assume 512, and serve mode additionally demands a canonical profile-filled cache. This PR threads the model's *own declared* expert geometry through those paths so pruned variants boot and run, while canonical behaviour stays byte-identical. ## Changes - **`tools/iq_pack.py`** — `N_EXPERT = 512` was hardcoded; the expert count is now derived from the router tensor's shape (`ffn_gate_inp.weight`, dim 1) with a uniformity check across layers. Blob sizes, offsets and the `native_experts.txt` header all follow it. - **`src/kernels/cpu/expert_layout.cpp`** — `expert_layout_load()` trusted the caller's geometry (a compile-time default) for the contiguity stride. It now parses `n_expert` from the pack header (v3 headers already record it) and uses that for both the stride and the stored layout, so `ArenaExpertSource`'s geometry check passes. - **`include/strata/artifact/gguf_reader.hpp`** — `Qwen4ExpGuard` validated every GGUF against canonical constants. `experts`/`experts_used` now default to 0 = presence-only; block count, hidden size and head counts are still hard-validated. - **`src/program/generate.cpp`** — the runtime `ModelGeometry` was default-constructed; it now reads `qwen4exp.expert_count` / `expert_used_count` from the `--native` shard (K too). The MTP drafter deliberately keeps the **canonical** geometry via a `static const ModelGeometry` — the draft layer is the original 512-expert head, and it sizes its own arena from it (speculation works unchanged; acceptance 74-85% measured). The serve-mode gate also no longer demands the profile-filled graph-hit path: an on-demand (`--expert-cache auto`) cache is a valid serve configuration. - **`src/core/layer.cpp`** — the native fused router stays canonical-only; non-canonical geometry falls back to the generic `router_top10` instead of erroring. - **`src/prefill/*`** — the `NE = 512` constant is replaced by the geometry's expert count in buffer sizing, host loops and residency indexing; `route_kernel` is templatized on registers-per-lane (512 → 16, 256 → 8, anything else takes the generic `router_top10`); `route()` gained an `n_expert` parameter; `moe_set_bytes()` likewise. ## Design notes - The model file is the authority on its own MoE shape; the compile-time defaults remain as fallbacks and for the canonical fast paths. - The draft/target split is intentional: pruning changes the target's routed experts but the shipped MTP draft is still the original head, so it must keep 512x10 sizing. - Profiles: the shipped `data/expert-profile.bin` is 48x512. For a pruned model we remapped the ranked (layer, expert) pairs onto the survivor indices using the release's published `tensor-allocation/*.rco-allocation.txt` (Section 2 records each survivor's original index) — 8000 → 5492 pairs survived remapping, trimmed to fit VRAM. A `make_profile.py` run against pruned-model routing would do this properly; happy to follow up with that as a companion tool. ## Testing (Ryzen 5 5600X, RTX 3090 24 GB, CUDA 13.1) - Pack built from the IQ1_M pruned release (shard 2 is byte-identical to the canonical n-gram table; symlinked, no re-download). - Pruned lane boots: 11 pool workers, 262144 ctx, KV int8, remapped profile (4500 slots), MTP draft over the canonical head. - **~3.5-hour real agentic coding session** through llama-swap (three.js scene build): peak 192k tokens, one omp compaction, zero engine errors, decode 27-42 tok/s (post-0.1.14 AVX2 i-quant kernels: 38.7 cold / 42.6 warm on fixed prompts vs 23.1/24.6 before this upgrade), draft acceptance 65-85%. - Canonical lanes unaffected: Q2_0 daily driver answers at its usual speed (53.6 tok/s spot check), IQ2_XS and vision configs untouched. ## Disclosure This PR was authored by GLM-5.3-Flash (ZCode agent); Pete (pjgmobile) tested it on real workloads overnight and submitted it.
No site
Links install, modelos, releases.