Pull requests / #1329

#1329 Fix batch MTP shared draft head type

open · @zetlaw · 0 评论 · 在 GitHub 查看

NVIDIA / CUDAModels & quantsWindows

描述

## Problem
On a single GPU, `--batch 2 --batch-mtp` with a populated `rt/draft_vocab.bin` starts with two slots but aborts the first batch admission with `mtp: unsupported native MMVQ GGML type`. The subset head is Q5_K, which the kernel already supports. A slot drafter shares the owner's subset-head weights and vocabulary but leaves `dhead_type_` at its default `-1`.

## Fix
Copy the subset-head GGML type when binding a shared MTP drafter. Document the existing model-backed two-slot regression command.

## Verification
- Built the patched engine in Release with CUDA 12.8 / SM89, native experts and K-quant MMQ enabled.
- On a real IQ3_S GGUF with a 106,299-token draft vocabulary, `tools/batch_test.py --batch 2 --n 2 --max-new 64 --extra "--batch-mtp --pcie-frac 0 --adapt-every 1000000"` produced **64 tokens identical to solo decoding in each slot**.
- A patched loopback HTTP server negotiated two slots; two simultaneous requests each completed 384 tokens without error. Its engine log captured batch windows containing both slots (`1,1,0,0`).
- The existing service was restored to the unmodified serial configuration after the isolated test. This PR does not change a deployed service.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。