Pull requests / #1010
#1010 batch: up to 3 MTP drafts per slot inside the 8-row window (2 lanes, code: 71.6 -> 84.3 tok/s)
closed · @xdbxdbx · 0 comentarios · En GitHub
BenchmarksNVIDIA / CUDAModels & quants
Descripción
Builds on #846 (its commit is the first one here); the change to review is the last commit. ## Problem With `MULTI_CONCURRENCY=TRUE`, #846 gives each batch slot one MTP draft: 2 rows per slot. The verifier's window holds 8 rows, so 2 lanes fill only half of it, and a lane alone in its group drafts 1 token where solo decoding drafts up to 3. ## Change - **Groups:** each slot gets a group of 1, 2 or 4 rows: its real token plus 0, 1 or 3 drafts, never more drafts than it holds or can still emit (max tokens, the context). Groups of 3 rows are not used, which bounds the layouts the graph cache sees. - **Row rotation:** slots enter the window in rotation while each fits with its smallest group; the rows left grow groups from 2 to 4. A slot left out goes first in the next window. When every slot fits, the start moves one slot, so the slot that grows first takes turns. - **Graph reuse:** chosen slots take their rows in ascending slot order, so a window's layout (and its captured graphs) depends only on which slots it holds and their group lengths. - **Acceptance:** a slot keeps the longest prefix of its group whose drafts the target picked. It stops before an end-of-reply token, or before a pick that leaves no room for the next token. - **Graph cache:** bounded at 32 layouts, the least recently used pair dropped first. A failed capture or instantiate drops both graphs of the pair, so a half-built pair is never launched. - **Errors:** a failed window prints `ERR` on stderr as well as stdout (the server logs only stderr). ## Validation RTX 4090, 64 GB RAM, Qwen3.8-Flash-Next IQ2_XS, `--resident-experts --kv int8 --kv-resident 32768 --spec 4`; single runs, 512 tokens per request. | | #846 (1 draft) | this PR (up to 3) | |---|---|---| | 2 lanes x 256K, code | 71.6 tok/s total | 84.3 tok/s total | | 2 lanes x 256K, prose | 75.2 tok/s total | 72.6 tok/s total | - **One lane alone:** 124-135 tok/s. - **4 lanes** (groups of 2 rows each when all four decode), 64K context: about 100 tok/s combined (46.75 ms per window, 4.65 rows on average).
En el sitio
Enlaces a install, modelos, releases.