Issues / #506

#506 RTX 5090 (sm_120): prefill ~3x slower on 0.1.34 than 0.1.31 (IQ3_S, single GPU)

closed · @gravitomagnetic · 1 Kommentare · Auf GitHub

BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quantsLinux

Beschreibung

# RTX 5090 (sm_120): prefill ~3x slower on 0.1.34 than 0.1.31 (IQ3_S, single GPU)

Reported by [gravitomagnetic](https://github.com/gravitomagnetic), on a Linux desktop.
Same machine and configuration as my community benchmark report (RTX 5090,
1M-context agent workload, 2026-10-02), so the 0.1.31 baseline below is from
that measured setup, not an estimate.

## Setup

| | |
|---|---|
| Model | Qwen3.8-Flash-Next **IQ3_S** (GSQ-RCO split GGUF, sha256 in my bench report) |
| GPU | single RTX 5090 (32 GiB), PCIe Gen5 x16, clocks not fixed, no clock-lock service |
| CPU / RAM | Ryzen 9 7950X, 122 GiB |
| OS / driver | Ubuntu 26.04.1, kernel 7.0.0-34, NVIDIA 595.91.07 |
| Engine | locally compiled, CUDA 13.4, `-DCMAKE_CUDA_ARCHITECTURES=120` |
| Config | `--max-context 1048576 --rope-scaling yarn --rope-scale 4 --kv int8 --kv-resident 32768 --expert-cache 7500 --mtp --spec 4 --spec-min-p 0.5`, prefill `auto`, low-RAM mode off |

## Results (engine-reported `timings` via the server API)

| engine | 32k cold prefill | large-prompt observation |
|---|---|---|
| 0.1.31 | **~3,284 tok/s** (median of 3 runs) | 512k: ~3,019 tok/s |
| 0.1.34 | **~1,086 tok/s** (2 runs) | a 345k prompt read in ~405 s (~850 tok/s) |

Decode is unaffected (130+ tok/s, expert-cache hit rate 97% on 0.1.34), so
this is specifically the prompt path. The 0.1.34 runs happened while the
server also carried live agent traffic; the prefill figures come from the
engine's own timings, which my bench report notes are unaffected by
co-tenant queueing.

## Suspect

`ae652b2` "prefill auto: chunks up to 32768 (was 8192)" (landed v0.1.32),
with follow-ups `3e31eea` (chunks above 8192 opt-in) and `4e0592b`
(STRATA_PREFILL_BF16X2 default off). I did not bisect 0.1.32/0.1.33 — the
regression is measured between 0.1.31 and 0.1.34 only.

The 0.1.31 startup log on this box shows:

```
strata serve: prompt chunk auto: 8192 tokens
strata serve: the prompt path borrows 2884 CUDA0 cache slots (5.46 GiB)
```

I did not capture the equivalent lines under 0.1.34 at the time (the box was
reverted to 0.1.31 the same evening to keep working). Reading #445, I know
these are the lines you look at: I have 0.1.34 and 0.1.35 built here and can
rerun the probe with the `streamed ring` / `prompt chunk auto` / `borrows`
lines plus a 3-run median on request.

## Notes

- 0.1.35's changelog shows no prefill-path changes, so I assume the
  regression is still present there; unverified on this hardware.
- Happy to test `--prefill auto:8192` on 0.1.34/0.1.35 as a confirmation if
  that pins the chunk size back.

Mehr auf der Site

Links zu Install, Modellen, Releases.