Pull requests / #54

#54 Support pruned expert variants (GSQ-RCO-Coder: 256 of 512 experts)

closed · merged 2026-09-28 · @pjgmobile · 0 コメント · GitHub で見る

BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsSecurity

本文

# Support pruned expert variants (GSQ-RCO-Coder: 256 of 512 experts)

## Motivation

ISTA-DASLab recently published `Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF` — the same
architecture with **half the routed experts pruned away per layer** (512 → 256,
`qwen4exp.expert_count = 256`, still 10 active). The engine currently specialises to the
canonical geometry in a number of places, so a pruned model fails at a different spot for
each one: the pack tool divides by a hardcoded 512, the pack loader validates against the
compile-time default, the runtime graph is default-constructed, the prefill kernels and
buffers assume 512, and serve mode additionally demands a canonical profile-filled cache.

This PR threads the model's *own declared* expert geometry through those paths so pruned
variants boot and run, while canonical behaviour stays byte-identical.

## Changes

- **`tools/iq_pack.py`** — `N_EXPERT = 512` was hardcoded; the expert count is now derived
  from the router tensor's shape (`ffn_gate_inp.weight`, dim 1) with a uniformity check
  across layers. Blob sizes, offsets and the `native_experts.txt` header all follow it.
- **`src/kernels/cpu/expert_layout.cpp`** — `expert_layout_load()` trusted the caller's
  geometry (a compile-time default) for the contiguity stride. It now parses
  `n_expert` from the pack header (v3 headers already record it) and uses that for both
  the stride and the stored layout, so `ArenaExpertSource`'s geometry check passes.
- **`include/strata/artifact/gguf_reader.hpp`** — `Qwen4ExpGuard` validated every GGUF
  against canonical constants. `experts`/`experts_used` now default to 0 = presence-only;
  block count, hidden size and head counts are still hard-validated.
- **`src/program/generate.cpp`** — the runtime `ModelGeometry` was default-constructed;
  it now reads `qwen4exp.expert_count` / `expert_used_count` from the `--native` shard
  (K too). The MTP drafter deliberately keeps the **canonical** geometry via a
  `static const ModelGeometry` — the draft layer is the original 512-expert head, and it
  sizes its own arena from it (speculation works unchanged; acceptance 74-85% measured).
  The serve-mode gate also no longer demands the profile-filled graph-hit path: an
  on-demand (`--expert-cache auto`) cache is a valid serve configuration.
- **`src/core/layer.cpp`** — the native fused router stays canonical-only; non-canonical
  geometry falls back to the generic `router_top10` instead of erroring.
- **`src/prefill/*`** — the `NE = 512` constant is replaced by the geometry's expert
  count in buffer sizing, host loops and residency indexing; `route_kernel` is
  templatized on registers-per-lane (512 → 16, 256 → 8, anything else takes the generic
  `router_top10`); `route()` gained an `n_expert` parameter; `moe_set_bytes()` likewise.

## Design notes

- The model file is the authority on its own MoE shape; the compile-time defaults remain
  as fallbacks and for the canonical fast paths.
- The draft/target split is intentional: pruning changes the target's routed experts but
  the shipped MTP draft is still the original head, so it must keep 512x10 sizing.
- Profiles: the shipped `data/expert-profile.bin` is 48x512. For a pruned model we
  remapped the ranked (layer, expert) pairs onto the survivor indices using the release's
  published `tensor-allocation/*.rco-allocation.txt` (Section 2 records each survivor's
  original index) — 8000 → 5492 pairs survived remapping, trimmed to fit VRAM. A
  `make_profile.py` run against pruned-model routing would do this properly; happy to
  follow up with that as a companion tool.

## Testing (Ryzen 5 5600X, RTX 3090 24 GB, CUDA 13.1)

- Pack built from the IQ1_M pruned release (shard 2 is byte-identical to the canonical
  n-gram table; symlinked, no re-download).
- Pruned lane boots: 11 pool workers, 262144 ctx, KV int8, remapped profile
  (4500 slots), MTP draft over the canonical head.
- **~3.5-hour real agentic coding session** through llama-swap (three.js scene build):
  peak 192k tokens, one omp compaction, zero engine errors, decode 27-42 tok/s
  (post-0.1.14 AVX2 i-quant kernels: 38.7 cold / 42.6 warm on fixed prompts vs 23.1/24.6
  before this upgrade), draft acceptance 65-85%.
- Canonical lanes unaffected: Q2_0 daily driver answers at its usual speed (53.6 tok/s
  spot check), IQ2_XS and vision configs untouched.

## Disclosure

This PR was authored by GLM-5.3-Flash (ZCode agent); Pete (pjgmobile) tested it on real
workloads overnight and submitted it.

関連リンク

インストール・モデル・リリースへの站内リンク。