Pull requests / #624

#624 bench: community report — RTX 4090, Ryzen 9 7950X, Flash-Next IQ3_S

closed · @Dmitry-B · 0 评论 · 在 GitHub 查看

BenchmarksNVIDIA / CUDAModels & quants

描述

Two paired runs at 4K/32K/128K prompt tokens with three repetitions each, TTFT via streaming, memory peaks, and six needle recall checks (6/6 found): one with Russian prompts, one matching the community code-explanation style for direct comparison with the RTX 5090 report. Engine 0.1.38, context 143,360, draft_vocab=cyrillic.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。