Pull requests / #233
#233 docs: add community benchmark guide and RTX 5090 results
closed · @hagope · 0 コメント · GitHub で見る
BenchmarksNVIDIA / CUDAModels & quantsDocumentation
本文
Strata has speed reports and a needle benchmark script, but no dedicated guide for users to contribute measurements from their own hardware. Add a community benchmark guide, a report template, and a README link, then use that format for an actual RTX 5090 report. The new report measures source-built Strata 0.1.29 (`d6708a4`) with the original Qwen3.8 Flash-Next IQ2_XS on an RTX 5090, Core Ultra 9 285K, and 64 GB RAM. Three serial runs each at 4,096, 32,768, and 128,000 prompt tokens produced median decode rates of 179.4, 175.7, and 165.0 tok/s. All nine speed requests reused zero prompt tokens and generated 256 tokens. The bundled needle test passed all six cases, with actual prompt lengths recorded. The report includes the benchmark script, exact configuration, raw per-run outputs and timings, engine log, sampled memory data, build details, verified GGUF hashes, and draft/pack provenance. It explicitly records swap growth and the reused prefix in the final needle case. These results cover one synthetic workload and configuration; they do not change headline performance claims or establish overall model quality. Validation: all nine speed requests and six recall checks completed; actual token counts and aggregate calculations were checked; report links and JSON were validated; Python scripts compile; `git diff --check` passes. Changes are documentation and benchmark artifacts only.
関連リンク
インストール・モデル・リリースへの站内リンク。