Issues / #1161
#1161 [Feature Request] Port ngram-mod & other ngram-* self-speculative techniques?
open · @Interpause · 3 comments · View on GitHub
Description
Opening this as more of a discussion rather than feature request. llama.cpp has a series of self-speculation techniques that use the convo history. Very useful when it comes to repeated one-shot coding, but more importantly agentic coding (most edit tools require you repeat what you are about to replace ad verbatim, and models tend to draft full snippets of the code during reasoning before writing out the exact same thing). See: <https://github.com/ggml-org/llama.cpp/blob/master/docs/speculative.md> Also it is possible to stack the different self-speculation techniques together which is what I did back in llama.cpp: ngram-mod + ngram-map-k4v + spec-dflash together. ngram-mod handles large copy and pastes, ngram-map-k4v handles smaller edits, and spec-dflash (dflash2) handles anything less than 5 tokens long.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.