Pull requests / #688

#688 Bench: RTX 5090 / Ryzen 9 9950X3D, IQ3_S, v0.1.34 vs v0.1.38 vs fused prompt kernels

closed · @CYoung83 · 0 コメント · GitHub で見る

BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsLinux

本文

Filing per #165. Native Linux (Ubuntu, CUDA 13.3), single RTX 5090 32 GB, Ryzen 9 9950X3D, 89 GiB RAM. IQ3_S pack, native experts, serve args in `strata-iq3_s.json`.

Three passes on identical hardware/config: v0.1.34 prebuilt, v0.1.38 built from source (`sm_120`), v0.1.38 with `STRATA_PF_FUSED=1`. 36 cells total (3 x 4 contexts x 3 repeats), every cell cold prefill (`0 reused`), no co-tenant traffic, engine/API TTFT agreement within ~1%.

Headline: v0.1.38 prefill +13-14% at 4k/32k over v0.1.34; fused kernels add another +4-11% at 4k and up on IQ3_S. Decode medians move but run-to-run spread is ±15% at fixed config, so no decode claim beyond medians; raw draft-acceptance counts per run are in `data/` for the #463 crowd.

One harness fix included: the shipped `strata-bench.py` 3.4 chars/token estimate overshoots on this corpus (~3.0 actual), so the 128k cell 400s. The patched copy parses the exact prompt token count from the 400 body and rescales until it fits. Worth upstreaming either as-is or as a probe-based ratio.

No needle runs or under-load power capture in this pass; happy to rerun if the shape needs it.

関連リンク

インストール・モデル・リリースへの站内リンク。