Pull requests / #882

#882 bench: community report, Flash-Next IQ2_XS vs a dense 27B as coding agents (RTX 4090 + 32 GB RAM)

closed · @T-Crypt · 0 评论 · 在 GitHub 查看

BenchmarksMulti-GPUNVIDIA / CUDADocumentation

描述

Community report: Flash-Next IQ2_XS on Strata 0.1.39 against a dense Qwen3.8-27B (NInfer engine) on one RTX 4090 with 32 GB RAM. Both ran as coding agents.

- Synthetic: Strata decodes faster (173-185 vs 127-140 tok/s) and reads a ~200K prompt 2x faster (3,833 vs 1,857 tok/s). Bug finding is a tie, and recall to 200K is 9/9 for both.
- Three real tickets, Claude Code subagents, three at once per engine, 150-minute cap: Strata 86/150, NInfer 112/150. On the two tickets both engines finished, Strata scored higher (86 vs 78).
- The Strata-relevant finding: with `conversation_cache_mib` 0, three interleaved agent conversations reused 32.4% of prompt tokens. A typical turn re-read about 107K tokens. That re-reading, not decode speed, decided the elapsed time. The report has the logged turns.

Data only: one new directory under `bench/results/`, no engine changes. Limits are in the README.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。