Pull requests / #453

#453 prefill: the draft layer's batched K/V for a ring too (KV streaming) - prefill +4.6% with --kv-resident

closed · @architectds · 0 comentários · No GitHub

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

Descrição

## What

With KV streaming (`--kv-resident`), the MTP drafter's K/V is a ring of its window: `kv_mode` 2, page p in slot
p % n_slots, over a host copy. `Prefill::draft_kv` (E-9) refused anything but `kv_mode` 0, so every prompt fell back to
the drafter's own pass (`MtpDrafter::prefill`), one graph per few rows over the window's cells. On a streamed session
that was about 4.5% of every long prompt.

- **The batched path now takes the ring.** It uses the same appends (`kv_append_q4` / `kv_append`) with the ring's
  page table and host copy, as a streamed main layer appends.
- **No two cells of a batch share a slot.**
  - The cells written are those the window can still reach (`first_needed`, as before), and the ring holds them.
  - A batch is also capped at the ring's slots minus one page, since a batch can straddle one extra page.
- **A/B switch.** `STRATA_MTP_BATCH_RING=0` restores the drafter's own pass.
- **Debug.** `STRATA_DRAFT_TIMING=1` prints the pass's time, synced.

## Measured

**Setup:**
- RTX 5070 Ti 16 GB on PCIe 3.0 x16 (X370), Ryzen 9 5900XT, DDR4-2133, Windows 11, CUDA 13.0, built on v0.1.34.
- Qwen3.8-Flash-Next IQ3_XXS, `--kv q4_0 --kv-resident 32768`, text-only args, `--prefill auto` (8,192-token chunks),
  `--spec 4 --mtp`.

**The draft layer's step after the prompt** (`STRATA_PREFILL_TIMING`'s "after each chunk"; the same prompts in both
arms, display off):

| Prompt tokens | Drafter's own pass | Batched (`STRATA_DRAFT_TIMING`) |
| ---: | ---: | --- |
| 6,949 | 204 ms | 17 ms: 6,942 cells in 16.4 ms |
| 14,297 | 421 ms | 34 ms: 8,192 + 6,098 cells in 19.5 + 14.6 ms |

**Prefill, end to end.**
- Only this switch differs: `STRATA_MTP_BATCH_RING=0` against the default, everything else as v0.1.34.
- The same three prompts in every run: a priming run first (not counted), then three rounds of on, off. Display off.

| Prompt tokens | Drafter's own pass (tok/s, 3 runs) | Batched (tok/s, 3 runs) | Change |
| ---: | ---: | ---: | ---: |
| 8,128 | 1,656 / 1,650 / 1,654 | 1,731 / 1,725 / 1,734 | +4.6% |
| 15,547 | 1,670 / 1,663 / 1,668 | 1,754 / 1,735 / 1,744 | +4.6% |
| 31,649 | 1,684 / 1,686 / 1,677 | 1,770 / 1,757 / 1,758 | +4.7% |

The saving grows with the prompt (median 215 / 407 / 873 ms). With a 32K-cell ring the drafter's window covers the
whole prompt, and its own pass cost about 27 ms per 1K prompt tokens.

**The drafts.** 128 greedy tokens after an 8K and a 16K prompt, two runs (the second with `STRATA_IQ_MT_MIN=1`):

| | Drafter's own pass | Batched |
| --- | --- | --- |
| Drafts accepted | 73/108, 64/101, 71/111, 73/107 (64-68%) | 75/117, 65/102, 69/113, 70/105 (61-67%) |
| Decode (tok/s) | 49.0, 59.9, 49.1, 58.4 | 46.6, 60.1, 47.5, 58.7 |

**Exactness.**
- The batched pass is not bitwise the drafter's own pass. As E-9 states for the non-ring case, its FP16/Q8 GEMMs
  round differently, so drafts can differ.
- The target's logits never depend on the drafts, but its greedy tokens can: a verify window's arithmetic depends on
  how many drafts it holds (#152).
- In the runs above the target's continuations parted after 2-35 tokens, also with `STRATA_IQ_MT_MIN=1`. That is the
  same thing a change of drafter or of `--spec` does.

## Not tested here

- An INT8 KV ring: the same `kv_append` with the ring's page table and host copy. Only `--kv q4_0` was run.
- HIP.
- Linux.
- A layer split: that path does not use `draft_kv`.

Developed with an AI coding assistant; all numbers measured on the machine above.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

No site

Links install, modelos, releases.