Pull requests / #708

#708 serve: --mtp is optional

closed · @merbanan · 0 评论 · 在 GitHub 查看

Server & APINVIDIA / CUDA

描述

`serve` required `--mtp DIR`, although `generate` already runs without it. Without the draft layer, `serve` now drafts with the suffix/prompt-lookup drafter only (or one token per round), the same way `generate` does. Every token is still verified against the model, so the output is the model's own; only speed changes.

This is useful on small cards: the draft layer and its head take roughly 0.7–1 GiB of VRAM, which goes to the expert cache instead. In an A/B on an RTX 2060 SUPER 8 GB, the cache grew from 1,000 to 1,678 slots.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。