Pull requests / #282

#282 prefill auto: chunks up to 32768 (was 8192), +15% on a 32K prompt

closed · @sergqwer · 0 commentaires · Sur GitHub

BenchmarksNVIDIA / CUDAModels & quantsWindows

Description

## Summary

With `--prefill auto` the largest chunk was 8192 tokens. Every chunk streams each expert it routes to again, so a
32K prompt crossed PCIe with nearly all experts of every layer four times. The auto list now starts at 32768 and
16384. It is still bounded by the lend rule (85-90% of the cache slots) and by `--max-context`. A request lends only
what its own prompt needs, so short prompts are unaffected.

Periodic prompt-cache checkpoints fall on chunk ends, so in a long prompt they now land every 32K rather than every
16K.

## Measured

Measured on current main (0.1.29, d6708a4) with ISTA-DASLab's IQ2_XS, RTX 5090 32 GB, Ryzen 9 9950X3D, 128 GB DDR5, Windows 11. KL is first-token KL divergence (`STRATA_DUMP_FIRST_LOGITS`, #276); two runs of main itself differ by 0 on 5 of 6 short prompts and by 0.00013 on the sixth. A 32K prompt, `--vram-reserve-mib 1500` (see #279), 3 runs each,
interleaved:

| | chunks | prompt speed | experts streamed |
| --- | ---: | ---: | ---: |
| main | 4 x 8192 | 5,513 / 5,687 / 5,671 tok/s | ~40,000 |
| this PR | 1 x 32768 | 6,231 / 6,548 / 6,615 tok/s | ~18,300 |

That is +15%. The 16 generated tokens were the same in 2 of the 3 pairs; one chunk instead of four changes the
summation order. An NVFP4 pack at 262K context gained more
(3,535 -> 5,201 tok/s): its experts are twice the size.

## Switch

`--prefill 8192` keeps the old chunk size.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

Sur le site

Liens install, modèles, releases.