Issues / #1729
#1729 0.1.41: STRATA_PREFILL_CPU_SHARE on by default halves decode throughput on a CPU without AVX-512 (i9-14900KF + RTX 5090 D, Windows), and auto decides sharing pays when it does not
open · @Simonqujian78 · 1 comentários · No GitHub
BenchmarksNVIDIA / CUDAWindows
Descrição
Summary On 0.1.41 my decode throughput dropped from a stable ~110 tok/s to ~46 tok/s (2.4x slower), and a cold 84k-token prompt read ran at 1,892 tok/s instead of ~5,066 tok/s. Both are the opposite of what the 0.1.41 release notes claim for this feature. Turning the new default off with STRATA_PREFILL_CPU_SHARE=0 restores both numbers (and higher). I believe auto misjudges how much share a CPU without AVX-512 can take, and that the resulting permanently-engaged CPU pool also slows down decode, which the feature was only measured on prompt time. [github-issue-0.1.41-cpu-share.md](https://github.com/user-attachments/files/33255929/github-issue-0.1.41-cpu-share.md)
No site
Links install, modelos, releases.