Pull requests / #1305
#1305 decode: --lookup-chain 1 and 2 do not change the window cap with the default --suffix-draft
open · @ZackO2o · 0 コメント · GitHub で見る
AMD / HIPNVIDIA / CUDAModels & quants
本文
## The change
Size the cap from the cap actually in force:
```cpp
if (o.lookup_chain > 0 && o.spec >= 2) {
if (o.mtp_max_t == 0) o.mtp_max_t = o.spec;
const int chain_base = o.spec;
o.spec = std::max(chain_base, std::min(std::max(chain_base, o.mtp_max_t) + o.lookup_chain, 8)); // kVerifyMaxT
}
```
Resulting caps with the default `--suffix-draft 3` and `--spec 4`:
| K | before | after |
| --: | --: | --: |
| 1 | 6 | **7** |
| 2 | 6 | **8** |
| 3 | 7 | **8** |
| 4+ | 8 | 8 |
The `kVerifyMaxT` = 8 ceiling holds, and behavior is unchanged where it should be — verified by evaluating both versions of the arithmetic:
- `--suffix-draft 0`: identical for K = 0, 2, 4, 7 (the `mtp_max_t == 0` branch still takes over).
- `--spec 6` (already near the ceiling, so the suffix block clamps at 8): identical for K = 1, 2, 3.
- Maximum reachable cap across K = 0..7 is 8, i.e. the ceiling is not crossed.
**This is a one-file, 5-line diff and I have not been able to build or run it.** Our box runs a self-built sm_70 engine from a different tree, and building this fork to validate a 5-line change is a larger operation than the change itself; I would rather flag that than claim a test I did not run. The arithmetic above is evaluated directly, and the "before" column is measured on 0.1.40 — that is the limit of what I can honestly assert.
If you would rather not take the code change, the alternative is one line of help text so the behavior is discoverable:
```
--lookup-chain K opt-in: after the MTP's drafts, add up to K prompt-lookup drafts that continue
them (the window grows to at most 8; default 0 = off)
NOTE: with the default --suffix-draft the cap is already --spec+2, so only
K > 2 raises it; K beyond the cap's headroom has no effect.
```
## The symptom
With the default `--suffix-draft 3`, **`--lookup-chain 1` and `--lookup-chain 2` produce a verify window identical to not passing the flag at all.** The option is documented as letting a chained lookup "lengthen one [window] by up to K", so a user setting `--lookup-chain 2` reasonably expects +2 headroom in the cap. They get +0, silently.
Measured on a 2× V100-PCIE-32GB box (sm_70, CUDA 12 source build, Qwen3.8-Flash-Next IQ3_S), reading the window cap back from the engine's own `/metrics` `engine` block (`spec` is the cap) after a restart per row:
| `--suffix-draft` | `--lookup-chain` | `spec` (window cap) | `mtp_max` |
| --: | --: | --: | --: |
| default | *(unset)* | 6 | 4 |
| default | **1** | **6** | 4 |
| default | **2** | **6** | 4 |
| default | 3 | 7 | 4 |
| default | 4 | 8 | 4 |
| 0 | 4 | 8 | 4 |
`--lookup-chain 1` and `2` are indistinguishable from unset. `3` and above do move the cap, but by less than K.
## The cause
The two blocks run in sequence in `generate.cpp` (~line 2089) and the second reads `mtp_max_t`, which the first has already written:
```cpp
// suffix drafter, on by default
if (o.suffix_draft > 0 && o.spec >= 2 && o.mtp_max_t == 0) {
o.mtp_max_t = o.spec; // 4
o.spec = std::min(o.spec + 2, 8); // 6
}
// --lookup-chain
if (o.lookup_chain > 0 && o.spec >= 2) {
if (o.mtp_max_t == 0) o.mtp_max_t = o.spec; // not taken: 4 != 0
o.spec = std::max(o.spec, std::min(o.mtp_max_t + o.lookup_chain, 8));
// = max(6, min(4 + K, 8)) <- K=1,2 give max(6, 5|6) = 6
}
```
Either the `mtp_max_t == 0` guard or the `mtp_max_t` term is stale by the time block 2 runs. With `--spec 4`, block 1 leaves `spec = 6`, so block 2's `min(mtp_max_t + K, 8) = min(4 + K, 8)` only exceeds 6 once `K >= 3` — and even then it is sizing against `mtp_max_t` (4), not against the `spec` (6) that was actually set.
The `--lookup-chain` line of `--help` says "the window grows to at most 8", which is true of the *ceiling* but not of the *increment*; there is no way for a user to discover that small K values are inert in the default configuration.
## Why this is worth fixing rather than documenting
The help text implies a monotonic relationship between K and the cap. A user sweeping K looking for the point where chaining starts to pay — which is the natural way to tune an opt-in flag — will see K=1 and K=2 do nothing and conclude, as we initially did, that the feature does not work. Our first write-up of the sweep reported `--lookup-chain 2` as a *negative* result (`-1.2% / -3.2%` vs default); reading this code afterwards is what showed the window never changed, so that row measured noise in a configuration where the flag could not have helped.
## What we measured, and what we did not
- The table above is from `/metrics` after a restart per row on 0.1.40, one box, one configuration. It is a reading of the cap, not a throughput claim.
- **We have not measured whether a larger cap actually helps on any workload.** The chained drafts are gated at runtime by the draft policy (`policy.chain(...)`) and by `T < S`, so a bigger cap is headroom, not a promise of accepted tokens. Our throughput sweeps of `--lookup-chain 2/4` on prose and code showed no gain, which is consistent with the mechanism documented in `--suffix-draft` (it pays on repeating context, which neither prompt is).
- Related community report on the same box: #1300. Workload note on `--lookup-chain`: #1301.
関連リンク
インストール・モデル・リリースへの站内リンク。