Pull requests / #882
#882 bench: community report, Flash-Next IQ2_XS vs a dense 27B as coding agents (RTX 4090 + 32 GB RAM)
closed · @T-Crypt · 0 comments · View on GitHub
BenchmarksMulti-GPUNVIDIA / CUDADocumentation
Description
Community report: Flash-Next IQ2_XS on Strata 0.1.39 against a dense Qwen3.8-27B (NInfer engine) on one RTX 4090 with 32 GB RAM. Both ran as coding agents. - Synthetic: Strata decodes faster (173-185 vs 127-140 tok/s) and reads a ~200K prompt 2x faster (3,833 vs 1,857 tok/s). Bug finding is a tie, and recall to 200K is 9/9 for both. - Three real tickets, Claude Code subagents, three at once per engine, 150-minute cap: Strata 86/150, NInfer 112/150. On the two tickets both engines finished, Strata scored higher (86 vs 78). - The Strata-relevant finding: with `conversation_cache_mib` 0, three interleaved agent conversations reused 32.4% of prompt tokens. A typical turn re-read about 107K tokens. That re-reading, not decode speed, decided the elapsed time. The report has the logged turns. Data only: one new directory under `bench/results/`, no engine changes. Limits are in the README.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.