Issues / #1252

#1252 DraftPolicy: a lookup window size priced during a slow stretch stays over-priced for the rest of the process

closed · @guthirry · 1 コメント · GitHub で見る

BenchmarksNVIDIA / CUDAWindows

本文

**Summary.** `DraftPolicy` updates a window size's cost (`cost_[t]`, an EMA of round times) only when that size is chosen. The `kProbes` re-probing covers only sizes measured fewer than three times. The policy lives for the whole process (`generate.cpp`: "learned over the whole process"). Suppose the wide lookup sizes are measured during a slow stretch, for example right after a prompt while the cache refills. Their cost stays high, so `choose()` keeps preferring short lookup windows. They are not measured again, and the speed loss carries over to every later request. The answers do not change; only decode speed does.

**What we saw.** Engine 82f46a8, 2026-10-06. Two GPUs: RTX 3090 (PCIe 4.0 x4) + RTX 3080 Ti, no P2P, `--peer-device 1`, UD-Q4_K_XL, `--spec 4`, default `--suffix-draft 3`. The workload was a 3,437-token prompt that asks for a 10 KB Python module back with two classes renamed (3,221 tokens, greedy). In the second build, the prompt pass split the experts between the GPUs differently. That was an experimental change to the prompt path only; decode was untouched.

| Server process | Answer | Decode | Tokens / window | ms / window | `suffix drafts` (server log) |
|---|---|---:|---:|---:|---|
| Build A, request 1 | identical | 97.6 tok/s | 5.42 | 55.6 | 456 windows, 2,245 of 2,249 accepted |
| Build A, request 2 | identical | 98.4 tok/s | 5.52 | 56.1 | 464 windows, 2,305 of 2,320 accepted |
| Build B, request 1 | identical | 83.2 tok/s | 3.91 | 47.0 | 480 windows, 1,433 of 1,436 accepted |
| Build B, request 2 | identical | 80.5 tok/s | 3.78 | 46.9 | 505 windows, 1,398 of 1,407 accepted |

In build B, the lookup windows stay as frequent and as accurate (99.7-99.8% accepted). They carry about 3.0 drafts instead of about 4.9. Each round is cheaper (47 vs 56 ms), but fewer tokens are committed per round, so decode is 15-18% slower. This persists across consecutive requests in the same process. In another process of a similar build, request 1 showed the same pattern (4.04 tokens per window, 79.7 tok/s) and request 2 recovered (5.51, 95.0 tok/s). That fits a learned, process-wide state rather than anything in the request.

**What we did not measure.** We did not log `cost_[]` itself, so the slow stretch that priced the wide sizes is inferred. The best candidate is the first decode rounds after a prompt that borrowed cache slots (`lent slots refilled ...`), when the second GPU is also busier. Printing `cost_[1..8]` and `cost_n_[]` at the end of each request would confirm or refute this quickly.

**Why it can lock in.** In `choose()`, a lookup of `k` drafts is taken only if `E(k) / cost_ms(k + 1)` beats the MTP's rate by `margin`. If `cost_[6]` was set high early, `best_t` lands on a smaller `k`, whose cost is measured often because the MTP uses those sizes. Size 6 is then rarely chosen, and its EMA (`kCostAlpha = 0.1`) never gets the samples it would need to come back down. `cost_n_[6] >= kProbes`, so the probe path does not re-try it either.

**Possible fixes.** These are suggestions only; we have not tested them.
- Re-probe a lookup size whose cost sample is old, for example after N rounds without a sample of that size, or with a small fixed exploration rate.
- Decay `cost_n_` (as `ok_`/`bad_` already decay) so the `kProbes` path re-opens after a while.
- Compare sizes by ratio to a size measured in the same period (for example `cost_[t] / cost_[t_mtp]`), so a global slow stretch does not price one size against the others.
- Skip `observe()` for rounds that overlap a known transient (the post-prompt refill), or reset the cost table per request.

The prompt-path change that exposed this is unrelated to the policy; its PR is https://github.com/Niko1221/Strata/pull/1251. On stock 82f46a8 the trigger would be any slow stretch that happens to coincide with the first samples of a wide size.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

関連リンク

インストール・モデル・リリースへの站内リンク。