Pull requests / #1381
#1381 mtp: the draft layer's experts at Q8_0 and BF16, alongside Q2_0
open · @gopinath87607 · 0 评论 · 在 GitHub 查看
BenchmarksNVIDIA / CUDAModels & quants
描述
The MTP draft layer's 512 routed experts were Q2_0 only: a format picked by measured draft acceptance, carrying a relative RMS error of 0.42 on the first experts. Higher precision buys better drafts, hence a higher accept rate, hence speed - never a different output, since the verify window decides every emitted token. This adds Q8_0 (2.49 GiB) and BF16 (4.69 GiB) as opt-in arms and leaves Q2_0 the default. Which arm a run uses is the runtime directory, not a flag: `--mtp rt-q8_0` or `--mtp rt-bf16`. The selection is an `experts.fmt` marker that tools/mtp_rt.py writes beside experts.bin, one line `<format> <gu_type> <d_type>`. No marker means Q2_0, so every `rt/` that exists today keeps taking exactly the path it always did. Engine: * iq_kernels.cu: Fmt<30> (BF16 weights against a q8_1 activation), and 30 in STRATA_GU_FMTS / STRATA_D_FMTS / STRATA_MMVQ_FMTS. Q8_0's grouped path at this shape is covered here for the first time as well. * `native_expert_supported` asks `embed_type_supported` rather than `is_iq`, which is the set dq_dispatch implements; `is_iq` is left alone, so nothing that gates on it moves. * mtp.cpp: one branch in load(), taken once. The per-expert stride is a function of the chosen layout rather than cpu::BLOB, and the activation buffer is carved at q8_1's 36 bytes per 32 values with Q2_0 reading its 34-byte prefix. A shared drafter now inherits the layout with the buffer it shares: without that a second session would stride a q8_0 buffer by the Q2_0 blob, silently. * The size of experts.bin is checked against the layout, because a mismatch does not crash - it drafts from the wrong bytes and is merely accepted less often. tools/mtp_pack.py gains a lossless bf16 arm (the checkpoint's own 16-bit words, no round trip, so its reconstruction error is exactly zero) and tools/mtp_rt.py writes the per-expert native blob: gate_up[e] then down[e]. Whole-tensor order would hand every expert the head of its neighbour. Measured on the 4-GPU rig (iq3_s, 256 greedy tokens): BF16 does not fit - the draft layer lands entirely on the last stage's 8 GB card and 4.84 GiB of experts leaves too few cache slots for the prompt path. Q8_0 runs at 48.9 tok/s against Q2_0's 49.5-49.9, with an accept rate indistinguishable within the engine's ~8-point run-to-run spread, at a 2007 -> 1134 slot cost on that card. So this is an arm to measure, not a win to claim. ## Title <!-- If Applicable, reference the GitHub issue --> Issue: Resolves # ## Summary <!-- Quick Summary of changes --> ## What changed <!-- Specifics on files changed, and what changes were made there --> ## Extra Notes <!-- Any extra notes, delete if there are none -->
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。