Pull requests / #1329
#1329 Fix batch MTP shared draft head type
open · @zetlaw · 0 commentaires · Sur GitHub
NVIDIA / CUDAModels & quantsWindows
Description
## Problem On a single GPU, `--batch 2 --batch-mtp` with a populated `rt/draft_vocab.bin` starts with two slots but aborts the first batch admission with `mtp: unsupported native MMVQ GGML type`. The subset head is Q5_K, which the kernel already supports. A slot drafter shares the owner's subset-head weights and vocabulary but leaves `dhead_type_` at its default `-1`. ## Fix Copy the subset-head GGML type when binding a shared MTP drafter. Document the existing model-backed two-slot regression command. ## Verification - Built the patched engine in Release with CUDA 12.8 / SM89, native experts and K-quant MMQ enabled. - On a real IQ3_S GGUF with a 106,299-token draft vocabulary, `tools/batch_test.py --batch 2 --n 2 --max-new 64 --extra "--batch-mtp --pcie-frac 0 --adapt-every 1000000"` produced **64 tokens identical to solo decoding in each slot**. - A patched loopback HTTP server negotiated two slots; two simultaneous requests each completed 384 tokens without error. Its engine log captured batch windows containing both slots (`1,1,0,0`). - The existing service was restored to the unmodified serial configuration after the isolated test. This PR does not change a deployed service.
Sur le site
Liens install, modèles, releases.