Pull requests / #1115
#1115 bench: community results, 2x RTX 5060 Ti 16 GB (PCIe gen3), Strata 0.1.34
closed · @Efs-O · 0 comentários · No GitHub
BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows
Descrição
Results-only submission following `docs/COMMUNITY_BENCHMARKS.md`. No engine changes. **Hardware:** 2× RTX 5060 Ti 16 GB on PCIe gen3 x8/x8, i9-9900KF, 128 GB DDR4-3200, Windows 10. An RTX 3060 runs the vision encoder only. **Model:** Flash-Next GSQ-RCO IQ3_S on Strata 0.1.34 (`1678de3`), with MTP, KV int8, and context 154,624. **What's in it** - **Speed** (guide format): short prompt, long cold prompt (~20.9K tokens), and long prompt with 16,384 reused tokens. 3 runs each, median and range, 256-token cap, rates from the engine's timing lines. Decode 56–63 tok/s; cold prefill ~1,108 tok/s. - **One agentic coding task** (semver library, 55 hidden tests, n=1): Flash-Next on Strata against Qwen3.8-27B Q6 on llama.cpp on the same machine. Paid hosted models on the same task are in a separate table, for reference only, so readers can place the local results. - `runs.json` has per-run data with engine log lines; `speed_bench.py` is the script. **Limitations:** n=1 for the agent task, which also ran on Strata 0.1.32. The prompt text for the speed runs is private code and is not included. The disconnect-after-a-long-tool-call stall we hit on 0.1.32 no longer reproduces on 0.1.34. The report notes this, and we are not opening an issue for it. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
No site
Links install, modelos, releases.