Issues / #577

#577 0.1.38: UD-Q4_K_XL prompts 15-40% slower on a 96 GB PC, from the unbuffered file tier (STRATA_UNBUFFERED_LOAD=0 restores them)

closed · @brenoperucchi · 2 comentarios · En GitHub

BenchmarksNVIDIA / CUDAModels & quantsWindows

Descripción

mise ~/.config/mise/config.toml tools: [email protected]
On 0.1.38 the Unsloth UD-Q4_K_XL model reads prompts noticeably slower than on 0.1.34 on this machine. The log shows
the file tier switching to unbuffered reads (#357/#362), and setting `STRATA_UNBUFFERED_LOAD=0` brings the speed back,
slightly above 0.1.34.

**Machine:** RTX 5090 32 GB, Ryzen 9 5950X (AVX2), 96 GB DDR4-3200, NVMe, Windows 11, driver 616.64. Release binaries
0.1.34 and 0.1.38.
**Model and config:** UD-Q4_K_XL (hand-made pack, GGUF read in place), the same configs as in #433:
`--resident-budget-gib 72|40 --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --max-context 32768
--kv int8`. With 72 GiB every expert the GPU cache does not hold (48.94 GiB) is resident in RAM; with 40 GiB the rest
come from the GGUF.
**Method:** as in #433, public prompts built from this repo, 1 warm-up + 3 measured runs, each run reading its whole
prompt fresh; engine `timings`, medians.

Prompt reading, tokens/s:

| Budget | Prompt | 0.1.34 | 0.1.38 | 0.1.38, `STRATA_UNBUFFERED_LOAD=0` |
|---|---|---:|---:|---:|
| 72 GiB | ~2.7K | 890 | 757 | 899 |
| 72 GiB | ~14.7K | 1,949 | 1,658 | 2,038 |
| 40 GiB | ~2.7K | 803 | 472 | 844 |
| 40 GiB | ~14.7K | 1,844 | 1,105 | 1,913 |

Decode is about the same in all of them. At start, 0.1.38 logs:

```
strata generate: the file tier reads unbuffered (2 of 48 probe reads from the file cache; 83.0 GiB available, 103.7 GiB of files)
```

and with the variable set:

```
strata generate: the file tier reads through the file cache (STRATA_UNBUFFERED_LOAD=0)
```

From `experts_unbuffered()` in `src/core/pinned.cu`, the decision is `room >= total_bytes` with `room = available -
arena_bytes - 4 GiB`, and `set_unbuffered()` gets the requested budget (`o.resident_budget`) as `arena_bytes`. Here that
is 83.0 - 72 - 4 = 7 GiB against 103.7 GiB of files (all four shards), so the reads go unbuffered. With 72 GiB the
requested budget is larger than what is actually resident (48.94 GiB), which makes `room` smaller than it is. Even at
40 GiB, where the experts outside the budget really do come from the files, the file cache was faster than unbuffered
reads on this PC with 96 GB.

The packs on 0.1.38 got faster on the same machine (Swift IQ3_XXS and IQ3_S, prompt reading +11-20% at 14.7K and
28.9K tokens), so this seems specific to the GGUF-in-place file tier. The numbers will also go into #433.

En el sitio

Enlaces a install, modelos, releases.