Pull requests / #693

#693 prefill: auto:16384 tries every 1024 tokens above 8192, and equal chunks - prompts 21-38% faster from 20K on a 16 GB card

closed · @architectds · 0 评论 · 在 GitHub 查看

BenchmarksSetup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindows

描述

**The problem.**
- A prompt chunk of 1,024 tokens or more (`stream_all_min()`) streams nearly every expert the GPU does not hold,
  whatever its length. So a long prompt's cost is its number of such chunks.
- With `--prefill auto:16384` (#282), auto tries only 16,384 and then falls back to 8,192. A card whose cache can lend
  13,312 but not 16,384 reads every prompt in 8,192-token chunks.
- On an RTX 5070 Ti 16 GB with IQ3_XXS (4,170 slots), that means 8 chunks for a 64K prompt where 5 would do.

**The change.**
1. **Finer auto sizes.** Above 8,192, auto now tries every 1,024 tokens up to the `auto:` limit, and takes the largest
   chunk whose buffers fit under the same lend cap as before.
   - With `auto:16384`, a card where 16,384 fits sees no change.
   - With `auto:32768`, the sizes between 16,384 and 32,768 become candidates too.
2. **Equal chunks.** A prompt segment reads in as few chunks as the chunk size allows, all the same size (rounded up
   to 256). For example, 20,036 tokens read as 3 × 6,912, not 2 × 8,192 + 3,652.
   - This costs no extra streaming, and it borrows no more slots than that count needs.
   - The old split stays when its last chunk would be under `stream_all_min()`. That chunk moves only the experts its
     own tokens route to, so 16,402 tokens still read as 2 × 8,192 + 18. Equal chunks there measured 22% slower.

`Prefill::stream_all_min_tokens()` exposes the threshold, so the rule follows `STRATA_PREFILL_STREAM_MIN`.

**Scope.**
- The chunk planning in `generate.cpp`: `request_chunk`, the auto size list in `plan_lend`, and the layer split's
  `pick`, which uses the same list.
- Plain `--prefill auto`, which setup writes, keeps its 8,192 limit; only (2) applies there.
- Fixed `--prefill N` is unchanged, apart from (2).

**Correctness.**
- Each build gave the same 16 greedy answer tokens in both of its runs.
- The 65K prompt read alone after the warm-up gives the same tokens on v0.1.38 and on this change.
- Where runs differ, they also differ within v0.1.38 itself. Its two runs answered the 100K prompt differently with
  4,171 and 4,163 expert slots: the answer can change with which experts the cache holds.
- Needles: 9 of 9 (32K, 128K and 262K at depths 10, 50 and 90) with `--prefill auto:16384`.

**Measured.**
- Setup: RTX 5070 Ti 16 GB on PCIe 3.0 x16, Ryzen 9 5900XT, 64 GB DDR4-2133, Windows 11, IQ3_XXS, 400K context, int8
  KV.
- Method: through the server, a fresh one per run, two rounds, prompts of this repository's code, the engine's own
  prompt speed.
- Details and every request: `bench/results/2026-10-03-prompt-chunks`.

| Prompt (tokens) | v0.1.38 | `--prefill auto` | `--prefill auto:16384` |
| --- | ---: | ---: | ---: |
| 9,010 | 1,511 | 1,517 (+0.3%) | 2,410 (+59%) |
| 16,402 | 2,289 | 2,303 (+0.6%) | 2,449 (+7%) |
| 20,036 | 2,101 | 2,148 (+2.2%) | 2,898 (+38%) |
| 26,700 | 2,143 | 2,172 (+1.4%) | 2,783 (+30%) |
| 32,716 | 2,570 | 2,583 (+0.5%) | 3,120 (+21%) |
| 40,013 | 2,546 | 2,547 (0.0%) | 3,088 (+21%) |
| 65,426 | 2,597 | 2,602 (+0.2%) | 3,392 (+31%) |
| 100,117 | 2,428 | 2,451 (+1.0%) | 3,275 (+35%) |

**Not measured:** AMD, multi-GPU, and `auto:32768` on cards where 16,384 fits but 32,768 does not.

**Overlap:** #547 changes the byte count behind `lend_slots`, on the lines just above `request_chunk`. The two edit
different lines, and with #547 a larger chunk can fit more often.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。