Issues / #1085

#1085 Performance regression after updating to v0.1.40 - RTX 5090

closed · @Predator75 · 1 comments · View on GitHub

BenchmarksSetup & installNVIDIA / CUDAModels & quantsWindows

Description

### Performance regression after updating to v0.1.40

Hi, I'm seeing a significant performance degradation after updating to Strata v0.1.40.

I'm running Strata on an **RTX 5090 32 GB** with **Qwen3.8-Flash-Next UD-IQ4_XS (Unsloth)**.

Before the update, with this model and essentially the same setup, I was seeing roughly **160-180 tok/s** during generation.

After updating to **v0.1.40**, generation performance dropped to roughly **50 tok/s**, sometimes even lower depending on the run.

My current relevant configuration is:

```text
Model: Qwen3.8-Flash-Next-UD-IQ4_XS
GPU: RTX 5090 32 GB
--expert-cache auto
--prefill auto
--spec 4
--max-context 131072
--kv int8
--kv-resident 32768
--resident-budget-gib 35
--vram-reserve-mib 700
--spec-min-p 0.5
parallel: 1
MTP enabled
```

The important point is that this is happening with **parallel = 1**, so the slowdown is not caused by running multiple concurrent slots.

I also noticed something interesting in the logs. In some runs Strata cannot allocate the requested 35 GiB resident budget and reduces the effective resident allocation to around **21.9 GiB** because of the available Windows commit. Those runs generate a very large amount of file reads and performance becomes much worse.

In another test where Strata managed to keep approximately **34 GiB resident in RAM** and reported **no file reads**, generation performance was around **85-92 tok/s**.

That is considerably better, but it is still well below the approximately **160-180 tok/s** I was getting before the update.

So I'm wondering whether something changed in v0.1.40 regarding:

- expert residency / expert-cache behavior
- Windows RAM/commit budget calculation
- PLE / SSD fallback
- MTP/speculative decoding
- automatic cache/residency decisions

I understand that **v0.1.40.1 uses the exact same engine as v0.1.40**, so I don't expect the 0.1.40.1 hotfix to change inference performance.

I'm happy to provide the complete startup/generation log or run specific A/B tests if useful.

Has anyone else observed a similar decode performance regression going from the previous engine to v0.1.40?

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.