Issues / #1729

#1729 0.1.41: STRATA_PREFILL_CPU_SHARE on by default halves decode throughput on a CPU without AVX-512 (i9-14900KF + RTX 5090 D, Windows), and auto decides sharing pays when it does not

open · @Simonqujian78 · 1 comentarios · En GitHub

BenchmarksNVIDIA / CUDAWindows

Descripción

Summary
On 0.1.41 my decode throughput dropped from a stable ~110 tok/s to ~46 tok/s (2.4x slower), and a cold
84k-token prompt read ran at 1,892 tok/s instead of ~5,066 tok/s. Both are the opposite of what the
0.1.41 release notes claim for this feature.
Turning the new default off with STRATA_PREFILL_CPU_SHARE=0 restores both numbers (and higher). I believe
auto misjudges how much share a CPU without AVX-512 can take, and that the resulting permanently-engaged
CPU pool also slows down decode, which the feature was only measured on prompt time.

[github-issue-0.1.41-cpu-share.md](https://github.com/user-attachments/files/33255929/github-issue-0.1.41-cpu-share.md)

En el sitio

Enlaces a install, modelos, releases.