Pull requests / #238

#238 Let the native router and the expert geometry take the count from the file

closed · @lukmanfauzie · 0 Kommentare · Auf GitHub

NVIDIA / CUDAModels & quants

Beschreibung

The Coder router change part of #151,  @Niko1221 asked to split out, onto v0.1.29 (`d6708a4`).
This is the one that needs its own byte gate on the Coder.

The Coder release is Qwen3.8-Flash-Next with half of its routed experts pruned away (256 of 512 per layer, ISTA-DASLab's RCO) and every other geometry number identical, so a hard-coded 512 is not a tuning assumption but a mis-indexing bug: `ffn_gate_inp.weight` is `[n_embd, 256]`, and reading 512 logits per token reads the NEXT token's logits as experts 256..511 - finite, correctly sized and wrong.

- `native_router_top10` takes `n_expert` (32..512 in multiples of 32, checked) and the device kernel derives `per_lane = n_expert / 32`, filling the unused slots with -INFINITY so the softmax max, sum and argmax ignore them exactly as if the tensor were shorter.  The unused slots must NOT enter the sum: adding -INFINITY to it makes the reciprocal -0 and every probability a NaN, which the NaN guard then replaces with -FLT_MAX - a router that picks arbitrary experts. One kernel therefore serves both counts with no second code path.
- The multi-token launch walks its rows at stride `n_expert` instead of 512, so the pruned pack's verify window is right too.
- `read_expert_count` reads the count off the GGUF itself (the router's last axis, cross-checked against `qwen4exp.expert_count`) and the loader refuses a pack and a model file that disagree, rather than routing a 256-expert pack through a 512-expert geometry.
- The drafter passes its own count to the router, as its gemv already does.  The drafter is loaded with its own geometry (512 - the draft block was never pruned), which is not the main model's when that is the Coder release, and the `ffn_gate_inp.weight` shape it reads is the drafter's.
- `tools/iq_pack.py` derives the count from the file's tensors so a pruned model packs correctly, and `tools/make_expert_profile.py` builds a profile for it (`--want-expert`, because a routing trace only bounds the count).
- `--dump-routing` now reports the records it wrote rather than the number of dispatches: it was `calls`, so a trace of zero bytes - a profile that reads as empty instead of broken - reported as a full one.

Only the single-token router contract is widened.  The verify window still takes the 512-expert kernel (`verify.cpp` guards `native_router_enabled() && NE == 512 && K == 10`, and `native_router_top10_multi` is `[n,512]` internally), and the prompt path needs neither - `prefill.cpp` already calls `route(...)` with the geometry's count.  Widening the multi kernel is a separate change with its own gate, so it is not smuggled in here.

`expert_layout.cpp` is left alone: it already reads the pruned count from a v3 pack's header, and the router takes its count from `ModelGeometry`, not from that field, so the stricter parse adds risk without changing what runs.

Mehr auf der Site

Links zu Install, Modellen, Releases.