Issues / #577
#577 0.1.38: UD-Q4_K_XL prompts 15-40% slower on a 96 GB PC, from the unbuffered file tier (STRATA_UNBUFFERED_LOAD=0 restores them)
closed · @brenoperucchi · 2 comments · View on GitHub
BenchmarksNVIDIA / CUDAModels & quantsWindows
Description
mise ~/.config/mise/config.toml tools: [email protected] On 0.1.38 the Unsloth UD-Q4_K_XL model reads prompts noticeably slower than on 0.1.34 on this machine. The log shows the file tier switching to unbuffered reads (#357/#362), and setting `STRATA_UNBUFFERED_LOAD=0` brings the speed back, slightly above 0.1.34. **Machine:** RTX 5090 32 GB, Ryzen 9 5950X (AVX2), 96 GB DDR4-3200, NVMe, Windows 11, driver 616.64. Release binaries 0.1.34 and 0.1.38. **Model and config:** UD-Q4_K_XL (hand-made pack, GGUF read in place), the same configs as in #433: `--resident-budget-gib 72|40 --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --max-context 32768 --kv int8`. With 72 GiB every expert the GPU cache does not hold (48.94 GiB) is resident in RAM; with 40 GiB the rest come from the GGUF. **Method:** as in #433, public prompts built from this repo, 1 warm-up + 3 measured runs, each run reading its whole prompt fresh; engine `timings`, medians. Prompt reading, tokens/s: | Budget | Prompt | 0.1.34 | 0.1.38 | 0.1.38, `STRATA_UNBUFFERED_LOAD=0` | |---|---|---:|---:|---:| | 72 GiB | ~2.7K | 890 | 757 | 899 | | 72 GiB | ~14.7K | 1,949 | 1,658 | 2,038 | | 40 GiB | ~2.7K | 803 | 472 | 844 | | 40 GiB | ~14.7K | 1,844 | 1,105 | 1,913 | Decode is about the same in all of them. At start, 0.1.38 logs: ``` strata generate: the file tier reads unbuffered (2 of 48 probe reads from the file cache; 83.0 GiB available, 103.7 GiB of files) ``` and with the variable set: ``` strata generate: the file tier reads through the file cache (STRATA_UNBUFFERED_LOAD=0) ``` From `experts_unbuffered()` in `src/core/pinned.cu`, the decision is `room >= total_bytes` with `room = available - arena_bytes - 4 GiB`, and `set_unbuffered()` gets the requested budget (`o.resident_budget`) as `arena_bytes`. Here that is 83.0 - 72 - 4 = 7 GiB against 103.7 GiB of files (all four shards), so the reads go unbuffered. With 72 GiB the requested budget is larger than what is actually resident (48.94 GiB), which makes `room` smaller than it is. Even at 40 GiB, where the experts outside the budget really do come from the files, the file cache was faster than unbuffered reads on this PC with 96 GB. The packs on 0.1.38 got faster on the same machine (Swift IQ3_XXS and IQ3_S, prompt reading +11-20% at 14.7K and 28.9K tokens), so this seems specific to the GGUF-in-place file tier. The numbers will also go into #433.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.