Issues / #506
#506 RTX 5090 (sm_120): prefill ~3x slower on 0.1.34 than 0.1.31 (IQ3_S, single GPU)
closed · @gravitomagnetic · 1 コメント · GitHub で見る
BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quantsLinux
本文
# RTX 5090 (sm_120): prefill ~3x slower on 0.1.34 than 0.1.31 (IQ3_S, single GPU) Reported by [gravitomagnetic](https://github.com/gravitomagnetic), on a Linux desktop. Same machine and configuration as my community benchmark report (RTX 5090, 1M-context agent workload, 2026-10-02), so the 0.1.31 baseline below is from that measured setup, not an estimate. ## Setup | | | |---|---| | Model | Qwen3.8-Flash-Next **IQ3_S** (GSQ-RCO split GGUF, sha256 in my bench report) | | GPU | single RTX 5090 (32 GiB), PCIe Gen5 x16, clocks not fixed, no clock-lock service | | CPU / RAM | Ryzen 9 7950X, 122 GiB | | OS / driver | Ubuntu 26.04.1, kernel 7.0.0-34, NVIDIA 595.91.07 | | Engine | locally compiled, CUDA 13.4, `-DCMAKE_CUDA_ARCHITECTURES=120` | | Config | `--max-context 1048576 --rope-scaling yarn --rope-scale 4 --kv int8 --kv-resident 32768 --expert-cache 7500 --mtp --spec 4 --spec-min-p 0.5`, prefill `auto`, low-RAM mode off | ## Results (engine-reported `timings` via the server API) | engine | 32k cold prefill | large-prompt observation | |---|---|---| | 0.1.31 | **~3,284 tok/s** (median of 3 runs) | 512k: ~3,019 tok/s | | 0.1.34 | **~1,086 tok/s** (2 runs) | a 345k prompt read in ~405 s (~850 tok/s) | Decode is unaffected (130+ tok/s, expert-cache hit rate 97% on 0.1.34), so this is specifically the prompt path. The 0.1.34 runs happened while the server also carried live agent traffic; the prefill figures come from the engine's own timings, which my bench report notes are unaffected by co-tenant queueing. ## Suspect `ae652b2` "prefill auto: chunks up to 32768 (was 8192)" (landed v0.1.32), with follow-ups `3e31eea` (chunks above 8192 opt-in) and `4e0592b` (STRATA_PREFILL_BF16X2 default off). I did not bisect 0.1.32/0.1.33 — the regression is measured between 0.1.31 and 0.1.34 only. The 0.1.31 startup log on this box shows: ``` strata serve: prompt chunk auto: 8192 tokens strata serve: the prompt path borrows 2884 CUDA0 cache slots (5.46 GiB) ``` I did not capture the equivalent lines under 0.1.34 at the time (the box was reverted to 0.1.31 the same evening to keep working). Reading #445, I know these are the lines you look at: I have 0.1.34 and 0.1.35 built here and can rerun the probe with the `streamed ring` / `prompt chunk auto` / `borrows` lines plus a 3-run median on request. ## Notes - 0.1.35's changelog shows no prefill-path changes, so I assume the regression is still present there; unverified on this hardware. - Happy to test `--prefill auto:8192` on 0.1.34/0.1.35 as a confirmation if that pins the chunk size back.
関連リンク
インストール・モデル・リリースへの站内リンク。