Pull requests / #1010

#1010 batch: up to 3 MTP drafts per slot inside the 8-row window (2 lanes, code: 71.6 -> 84.3 tok/s)

closed · @xdbxdbx · 0 コメント · GitHub で見る

BenchmarksNVIDIA / CUDAModels & quants

本文

Builds on #846 (its commit is the first one here); the change to review is the last commit.

## Problem

With `MULTI_CONCURRENCY=TRUE`, #846 gives each batch slot one MTP draft: 2 rows per slot. The verifier's window holds 8 rows, so 2 lanes fill only half of it, and a lane alone in its group drafts 1 token where solo decoding drafts up to 3.

## Change

- **Groups:** each slot gets a group of 1, 2 or 4 rows: its real token plus 0, 1 or 3 drafts, never more drafts than it holds or can still emit (max tokens, the context). Groups of 3 rows are not used, which bounds the layouts the graph cache sees.
- **Row rotation:** slots enter the window in rotation while each fits with its smallest group; the rows left grow groups from 2 to 4. A slot left out goes first in the next window. When every slot fits, the start moves one slot, so the slot that grows first takes turns.
- **Graph reuse:** chosen slots take their rows in ascending slot order, so a window's layout (and its captured graphs) depends only on which slots it holds and their group lengths.
- **Acceptance:** a slot keeps the longest prefix of its group whose drafts the target picked. It stops before an end-of-reply token, or before a pick that leaves no room for the next token.
- **Graph cache:** bounded at 32 layouts, the least recently used pair dropped first. A failed capture or instantiate drops both graphs of the pair, so a half-built pair is never launched.
- **Errors:** a failed window prints `ERR` on stderr as well as stdout (the server logs only stderr).

## Validation

RTX 4090, 64 GB RAM, Qwen3.8-Flash-Next IQ2_XS, `--resident-experts --kv int8 --kv-resident 32768 --spec 4`; single runs, 512 tokens per request.

| | #846 (1 draft) | this PR (up to 3) |
|---|---|---|
| 2 lanes x 256K, code | 71.6 tok/s total | 84.3 tok/s total |
| 2 lanes x 256K, prose | 75.2 tok/s total | 72.6 tok/s total |

- **One lane alone:** 124-135 tok/s.
- **4 lanes** (groups of 2 rows each when all four decode), 64K context: about 100 tok/s combined (46.75 ms per window, 4.65 rows on average).

関連リンク

インストール・モデル・リリースへの站内リンク。