Issues / #1085
#1085 Performance regression after updating to v0.1.40 - RTX 5090
closed · @Predator75 · 1 Kommentare · Auf GitHub
BenchmarksSetup & installNVIDIA / CUDAModels & quantsWindows
Beschreibung
### Performance regression after updating to v0.1.40 Hi, I'm seeing a significant performance degradation after updating to Strata v0.1.40. I'm running Strata on an **RTX 5090 32 GB** with **Qwen3.8-Flash-Next UD-IQ4_XS (Unsloth)**. Before the update, with this model and essentially the same setup, I was seeing roughly **160-180 tok/s** during generation. After updating to **v0.1.40**, generation performance dropped to roughly **50 tok/s**, sometimes even lower depending on the run. My current relevant configuration is: ```text Model: Qwen3.8-Flash-Next-UD-IQ4_XS GPU: RTX 5090 32 GB --expert-cache auto --prefill auto --spec 4 --max-context 131072 --kv int8 --kv-resident 32768 --resident-budget-gib 35 --vram-reserve-mib 700 --spec-min-p 0.5 parallel: 1 MTP enabled ``` The important point is that this is happening with **parallel = 1**, so the slowdown is not caused by running multiple concurrent slots. I also noticed something interesting in the logs. In some runs Strata cannot allocate the requested 35 GiB resident budget and reduces the effective resident allocation to around **21.9 GiB** because of the available Windows commit. Those runs generate a very large amount of file reads and performance becomes much worse. In another test where Strata managed to keep approximately **34 GiB resident in RAM** and reported **no file reads**, generation performance was around **85-92 tok/s**. That is considerably better, but it is still well below the approximately **160-180 tok/s** I was getting before the update. So I'm wondering whether something changed in v0.1.40 regarding: - expert residency / expert-cache behavior - Windows RAM/commit budget calculation - PLE / SSD fallback - MTP/speculative decoding - automatic cache/residency decisions I understand that **v0.1.40.1 uses the exact same engine as v0.1.40**, so I don't expect the 0.1.40.1 hotfix to change inference performance. I'm happy to provide the complete startup/generation log or run specific A/B tests if useful. Has anyone else observed a similar decode performance regression going from the previous engine to v0.1.40?
Mehr auf der Site
Links zu Install, Modellen, Releases.