Pull requests / #1268

#1268 --batch-mtp: the first slot admission ends the engine ("unsupported native MMVQ GGML type")

closed · @crazyaimachine · 0 评论 · 在 GitHub 查看

Setup & installNVIDIA / CUDA

描述

**Problem.** With `--batch-mtp`, the engine exits at the first request that goes into a batch slot:

```
strata batch: MTP admission for slot 0 failed: mtp: unsupported native MMVQ GGML type
```

The slot drafters share the main drafter's weights and draft head (`MtpDrafter::bind(..., shared)`). `bind()` copies the draft head subset (`dhead_`, `dvocab_`, `n_dvocab_`) but not its ggml type. `dhead_type_` stays -1, and `native_mmvq` refuses it. This hits every `--batch-mtp` run that has a draft vocabulary, which is the default.

**Fix.** One commit in `src/core/mtp.cpp`:

- `bind()` with a shared drafter also copies `dhead_type_`.
- `load()` with shared weights also takes the `--mtp-q4` Q4_0 copies (`dense4_`, `q4_off_`, `q4_`, `q4_head_`). They would be missing in the same way.
- `dense4_` is freed only by the drafter that owns the weights.

**Checked.** Setup: 1× TITAN RTX (Turing), UD-Q4_K_XL with the BF16 PLE table, `--batch 4 --batch-mtp --spec 4`, MTP runtime with draft vocabulary.

- Before the fix: the engine exits at the first `BGEN`.
- After the fix: `tools/batch_test.py --batch 4 --n 4 --max-new 96 --mt-min 1 --extra "--batch-mtp --pcie-frac 0 --adapt-every 1000000"` reports all 4 slots IDENTICAL to their solo tokens.

Side note from the same machine: I also tried `--batch-mtp` on a 2-GPU layer split, without `--batch-groups`. It works and gives tokens identical to solo, but it needs two more changes, and on Turing it was slower than the pipelined batch without MTP. So it is not part of this PR. I can open it separately if it is of interest.

Written by my agent, tested on my machine.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。