Pull requests / #1329

#1329 Fix batch MTP shared draft head type

open · @zetlaw · 0 コメント · GitHub で見る

NVIDIA / CUDAModels & quantsWindows

本文

## Problem
On a single GPU, `--batch 2 --batch-mtp` with a populated `rt/draft_vocab.bin` starts with two slots but aborts the first batch admission with `mtp: unsupported native MMVQ GGML type`. The subset head is Q5_K, which the kernel already supports. A slot drafter shares the owner's subset-head weights and vocabulary but leaves `dhead_type_` at its default `-1`.

## Fix
Copy the subset-head GGML type when binding a shared MTP drafter. Document the existing model-backed two-slot regression command.

## Verification
- Built the patched engine in Release with CUDA 12.8 / SM89, native experts and K-quant MMQ enabled.
- On a real IQ3_S GGUF with a 106,299-token draft vocabulary, `tools/batch_test.py --batch 2 --n 2 --max-new 64 --extra "--batch-mtp --pcie-frac 0 --adapt-every 1000000"` produced **64 tokens identical to solo decoding in each slot**.
- A patched loopback HTTP server negotiated two slots; two simultaneous requests each completed 384 tokens without error. Its engine log captured batch windows containing both slots (`1,1,0,0`).
- The existing service was restored to the unmodified serial configuration after the isolated test. This PR does not change a deployed service.

関連リンク

インストール・モデル・リリースへの站内リンク。