Issues / #1301
#1301 docs: `--lookup-chain` — note the workload it targets (repeating context), since the default suffix path already covers non-repeating text
open · @ZackO2o · 0 コメント · GitHub で見る
BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
本文
**What this is:** a short measurement note on `--lookup-chain`, plus a documentation observation. Not a bug report.
**Hardware:** 2× Tesla V100-PCIE-32GB (sm_70, CUDA 12 source build), Qwen3.8-Flash-Next IQ3_S, 512K, `--layer-split 20`, `--spec 4 --spec-min-p 0.70`. Same box as the community report in #1300.
## What we measured
We swept three opt-in settings, 12 runs each, first 2 discarded, medians, arms interleaved within one service lifetime:
| Setting | prose | code | vs default |
| --: | --: | --: | --: |
| **default** | **77.0** | **103.1** | — |
| `--lookup-chain 2` | 76.1 | 99.8 | -1.2% / -3.2% |
| `--lookup-chain 4` | 73.7 | 100.8 | -4.3% / -2.2% |
| `--host-core last` | 76.3 | 101.2 | -0.9% / -1.8% |
`--host-core last` we understand — its help says the win comes from Windows sending a GPU's interrupts to one logical processor, which is a Windows-only condition; on Linux there is nothing to fix. Correctly documented, no action needed.
**`--lookup-chain` is the one worth a sentence in the help.** Reading the source (`generate.cpp:2089`), the default path already enables prompt lookup:
```cpp
// Prompt lookup (the suffix drafter, on by default): the MTP keeps its --spec windows and a lookup window may be
// up to 2 tokens longer; the draft policy (strata/spec/draft_policy.hpp) takes one only where it pays. Code
// edits +6-11%, ordinary text unchanged (bench/results/2026-09-27-spec). --suffix-draft 0 turns it off.
if (o.suffix_draft > 0 && o.spec >= 2 && o.mtp_max_t == 0) {
o.mtp_max_t = o.spec;
o.spec = std::min(o.spec + 2, 8);
}
```
So the default already takes the suffix path when the draft policy says it pays, and the +6–11% case documented there is **code edits**. `--lookup-chain K` is specifically "add up to K prompt-lookup drafts *that continue the MTP's drafts*", which needs the context to actually repeat itself — our prose and code prompts do not, so a null-to-negative result is the expected outcome rather than a surprising one.
## The suggestion
The help entry describes the mechanism precisely but not the workload it targets:
```
--lookup-chain K opt-in: after the MTP's drafts, add up to K prompt-lookup drafts that continue
them (the window grows to at most 8; default 0 = off)
```
It is opt-in and off by default, so nobody loses anything by it — but "opt-in" invites trying it, and the setting that decides the outcome (does your context repeat?) is not mentioned but *is* mentioned on the sibling `--suffix-draft` line ("when it pays"). A phrase in the same spirit on the `--lookup-chain` line — **"pays when the context repeats (templated output, fixed-format turns, code with repeated blocks)"** — would save the next person the sweep we just did.
## What we did not establish
- We have **not** tested `--lookup-chain` on a templated workload, so this is not a claim that it never helps — only that on non-repeating prose/code it does not, and that that is consistent with the mechanism.
- Our sweep ran in one batch each; the numbers above are the within-batch comparison, and we would not quote the absolute values across batches (this box has a documented content-dependence in decode rate — see #1300).
関連リンク
インストール・モデル・リリースへの站内リンク。