Pull requests / #1381

#1381 mtp: the draft layer's experts at Q8_0 and BF16, alongside Q2_0

open · @gopinath87607 · 0 comentarios · En GitHub

BenchmarksNVIDIA / CUDAModels & quants

Descripción

The MTP draft layer's 512 routed experts were Q2_0 only: a format picked by measured draft acceptance, carrying a relative RMS error of 0.42 on the first experts.  Higher precision buys better drafts, hence a higher accept rate, hence speed - never a different output, since the verify window decides every emitted token.  This adds Q8_0 (2.49 GiB) and BF16 (4.69 GiB) as opt-in arms and leaves Q2_0 the default.

Which arm a run uses is the runtime directory, not a flag: `--mtp rt-q8_0` or `--mtp rt-bf16`.  The selection is an `experts.fmt` marker that tools/mtp_rt.py writes beside experts.bin, one line `<format> <gu_type> <d_type>`.  No marker means Q2_0, so every `rt/` that exists today keeps taking exactly the path it always did.

Engine:
  * iq_kernels.cu: Fmt<30> (BF16 weights against a q8_1 activation), and 30 in STRATA_GU_FMTS / STRATA_D_FMTS / STRATA_MMVQ_FMTS.  Q8_0's grouped path at this shape is covered here for the first time as well.
  * `native_expert_supported` asks `embed_type_supported` rather than `is_iq`, which is the set dq_dispatch implements; `is_iq` is left alone, so nothing that gates on it moves.
  * mtp.cpp: one branch in load(), taken once.  The per-expert stride is a function of the chosen layout rather than cpu::BLOB, and the activation buffer is carved at q8_1's 36 bytes per 32 values with Q2_0 reading its 34-byte prefix.  A shared drafter now inherits the layout with the buffer it shares: without that a second session would stride a q8_0 buffer by the Q2_0 blob, silently.
  * The size of experts.bin is checked against the layout, because a mismatch does not crash - it drafts from the wrong bytes and is merely accepted less often.

tools/mtp_pack.py gains a lossless bf16 arm (the checkpoint's own 16-bit words, no round trip, so its reconstruction error is exactly zero) and tools/mtp_rt.py writes the per-expert native blob: gate_up[e] then down[e].  Whole-tensor order would hand every expert the head of its neighbour.

Measured on the 4-GPU rig (iq3_s, 256 greedy tokens): BF16 does not fit - the draft layer lands entirely on the last stage's 8 GB card and 4.84 GiB of experts leaves too few cache slots for the prompt path.  Q8_0 runs at 48.9 tok/s against Q2_0's 49.5-49.9, with an accept rate indistinguishable within the engine's ~8-point run-to-run spread, at a 2007 -> 1134 slot cost on that card.  So this is an arm to measure, not a win to claim.

## Title
<!-- If Applicable, reference the GitHub issue -->
Issue: Resolves # 

## Summary
<!-- Quick Summary of changes -->

## What changed
<!-- Specifics on files changed, and what changes were made there -->

## Extra Notes
<!-- Any extra notes, delete if there are none -->

En el sitio

Enlaces a install, modelos, releases.