Pull requests / #708

#708 serve: --mtp is optional

closed · @merbanan · 0 コメント · GitHub で見る

Server & APINVIDIA / CUDA

本文

`serve` required `--mtp DIR`, although `generate` already runs without it. Without the draft layer, `serve` now drafts with the suffix/prompt-lookup drafter only (or one token per round), the same way `generate` does. Every token is still verified against the model, so the output is the model's own; only speed changes.

This is useful on small cards: the draft layer and its head take roughly 0.7–1 GiB of VRAM, which goes to the expert cache instead. In an A/B on an RTX 2060 SUPER 8 GB, the cache grew from 1,000 to 1,678 slots.

関連リンク

インストール・モデル・リリースへの站内リンク。