Pull requests / #453
#453 prefill: the draft layer's batched K/V for a ring too (KV streaming) - prefill +4.6% with --kv-resident
closed · @architectds · 0 コメント · GitHub で見る
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux
本文
## What With KV streaming (`--kv-resident`), the MTP drafter's K/V is a ring of its window: `kv_mode` 2, page p in slot p % n_slots, over a host copy. `Prefill::draft_kv` (E-9) refused anything but `kv_mode` 0, so every prompt fell back to the drafter's own pass (`MtpDrafter::prefill`), one graph per few rows over the window's cells. On a streamed session that was about 4.5% of every long prompt. - **The batched path now takes the ring.** It uses the same appends (`kv_append_q4` / `kv_append`) with the ring's page table and host copy, as a streamed main layer appends. - **No two cells of a batch share a slot.** - The cells written are those the window can still reach (`first_needed`, as before), and the ring holds them. - A batch is also capped at the ring's slots minus one page, since a batch can straddle one extra page. - **A/B switch.** `STRATA_MTP_BATCH_RING=0` restores the drafter's own pass. - **Debug.** `STRATA_DRAFT_TIMING=1` prints the pass's time, synced. ## Measured **Setup:** - RTX 5070 Ti 16 GB on PCIe 3.0 x16 (X370), Ryzen 9 5900XT, DDR4-2133, Windows 11, CUDA 13.0, built on v0.1.34. - Qwen3.8-Flash-Next IQ3_XXS, `--kv q4_0 --kv-resident 32768`, text-only args, `--prefill auto` (8,192-token chunks), `--spec 4 --mtp`. **The draft layer's step after the prompt** (`STRATA_PREFILL_TIMING`'s "after each chunk"; the same prompts in both arms, display off): | Prompt tokens | Drafter's own pass | Batched (`STRATA_DRAFT_TIMING`) | | ---: | ---: | --- | | 6,949 | 204 ms | 17 ms: 6,942 cells in 16.4 ms | | 14,297 | 421 ms | 34 ms: 8,192 + 6,098 cells in 19.5 + 14.6 ms | **Prefill, end to end.** - Only this switch differs: `STRATA_MTP_BATCH_RING=0` against the default, everything else as v0.1.34. - The same three prompts in every run: a priming run first (not counted), then three rounds of on, off. Display off. | Prompt tokens | Drafter's own pass (tok/s, 3 runs) | Batched (tok/s, 3 runs) | Change | | ---: | ---: | ---: | ---: | | 8,128 | 1,656 / 1,650 / 1,654 | 1,731 / 1,725 / 1,734 | +4.6% | | 15,547 | 1,670 / 1,663 / 1,668 | 1,754 / 1,735 / 1,744 | +4.6% | | 31,649 | 1,684 / 1,686 / 1,677 | 1,770 / 1,757 / 1,758 | +4.7% | The saving grows with the prompt (median 215 / 407 / 873 ms). With a 32K-cell ring the drafter's window covers the whole prompt, and its own pass cost about 27 ms per 1K prompt tokens. **The drafts.** 128 greedy tokens after an 8K and a 16K prompt, two runs (the second with `STRATA_IQ_MT_MIN=1`): | | Drafter's own pass | Batched | | --- | --- | --- | | Drafts accepted | 73/108, 64/101, 71/111, 73/107 (64-68%) | 75/117, 65/102, 69/113, 70/105 (61-67%) | | Decode (tok/s) | 49.0, 59.9, 49.1, 58.4 | 46.6, 60.1, 47.5, 58.7 | **Exactness.** - The batched pass is not bitwise the drafter's own pass. As E-9 states for the non-ring case, its FP16/Q8 GEMMs round differently, so drafts can differ. - The target's logits never depend on the drafts, but its greedy tokens can: a verify window's arithmetic depends on how many drafts it holds (#152). - In the runs above the target's continuations parted after 2-35 tokens, also with `STRATA_IQ_MT_MIN=1`. That is the same thing a change of drafter or of `--spec` does. ## Not tested here - An INT8 KV ring: the same `kv_append` with the ring's page table and host copy. Only `--kv q4_0` was run. - HIP. - Linux. - A layer split: that path does not use `draft_kv`. Developed with an AI coding assistant; all numbers measured on the machine above. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
関連リンク
インストール・モデル・リリースへの站内リンク。