Pull requests / #1220
#1220 fix #1121: preserve shared draft-head metadata in batch MTP slots
closed · @x00r · 0 コメント · GitHub で見る
AMD / HIPNVIDIA / CUDAModels & quantsWindows
本文
# Preserve shared draft-head metadata in batch MTP slots ## Summary - Copy the source drafter's `dhead_type_` and `dvocab_host_` when a batch slot binds shared draft-head buffers. - Preserve existing ownership: the source drafter owns the device buffers and outlives the slots. - Add a small synthetic GPU regression test and opt-in bind diagnostics (`STRATA_DEBUG_MTP=1`, off by default). ## Problem With `--batch 2 --batch-mtp` and a draft vocabulary subset, batch admission repeatedly reports: ```text strata batch: MTP admission for slot 0 failed: mtp: unsupported native MMVQ GGML type ``` `MtpDrafter::bind()` shares `dhead_`, `dvocab_`, and `n_dvocab_`, but leaves the slot's `dhead_type_` at its default `-1`. The draft-output path passes that invalid format to `native_mmvq()`, and the admission error path exits the entire engine. Both slots can be affected while solo requests continue to work. The same sharing branch also omits the host vocabulary map used by `record_top2()` diagnostics. ## Fix Copy both metadata fields alongside the borrowed buffers: ```cpp dhead_type_ = shared->dhead_type_; dvocab_host_ = shared->dvocab_host_; ``` The format must come from the source drafter, not `head->type()`: a Q4-converted draft subset can differ from the main head. Keep `owns_draft_head_ = false`. No changes to model files, quantization formats, control vectors, or experimental speed-projection settings are needed. ## Tests and evidence On Windows, MSVC, CUDA 12.4, RTX 3090 / sm_86: - Engine and test build completed successfully. - Four targeted CTest tests passed: `mtp_shared_head_test`, `vmm_test`, `file_expert_source_test`, and `expert_profile_save_test`. - Removing only the format copy and rebuilding failed the regression test with `shared draft head type is -1, expected 13`. - Restoring the format copy but removing the token-map copy independently failed with `shared head lost the host token map (top2 diagnostics)`. - Restoring both copies and rebuilding passed all four tests again. The new test calls the real binding method and covers Q5_K, IQ4_XS, a Q4-converted subset, the no-subset fallback, both slots, incompatible-head rejection, and source-buffer survival after slot destruction. It does not launch draft-output MMVQ or run full-model concurrent generation. The full suite was not run. ## Scope / review notes - This patch includes only `src/core/mtp.cpp`, `include/strata/core/mtp.hpp`, `tests/core/mtp_shared_head_test.cpp`, and its CMake registration. - Local CUDA 12.4 validation also required a separate VMM driver-query compatibility fix; it is intentionally excluded from this PR patch.
関連リンク
インストール・モデル・リリースへの站内リンク。