Pull requests / #466

#466 bench: community report — RTX 5090, IQ3_S at 1M context (YaRN), agent-style workload

closed · @gravitomagnetic · 0 コメント · GitHub で見る

BenchmarksServer & APINVIDIA / CUDAModels & quantsDocumentation

本文

Results-only community benchmark per docs/COMMUNITY_BENCHMARKS.md. No engine changes.

**Hardware:** RTX 5090 32 GB (Gen5 x16), Ryzen 9 7950X, 122 GB RAM, Ubuntu 26.04.1, driver 595.91.07, source commit 9259cad (v0.1.31), CUDA 13.4.2.

**Model/config:** Qwen3.8-Flash-Next IQ3_S (split GGUF, sha256 recorded), --max-context 1048576 --rope-scaling yarn --rope-scale 4, --kv int8 --kv-resident 32768, --expert-cache 7500 (auto-fills to 9,804 slots), MTP --spec 4 --spec-min-p 0.5, vision on (no images in requests).

**What's measured:**
- Cold prefill ~2,500 tok/s at 4k, ~3,300 at 32k, ~3,400 at 128k, ~3,000 at 512k (engine-reported timings).
- Decode 94-105 tok/s with reasoning enabled (temp 1.0 / top_p 0.95 per the HF card); 71-87 tok/s greedy instruct mode.
- Prefix reuse: 500,553 of 500,558 tokens reused in 36 ms on warm repeat.
- Expert cache hit rate 96-98% during 512k runs.
- Needle recall via tools/needle_bench.py: 5/5 found at depth 50%, up to a 1,048,265-token prompt.

**Honest caveats (Note A in the README):** an apparent ~65 s per-request stall in early drafts was disproven by a controlled probe and re-attributed to serial scheduling behind a co-tenant 208k-token prefill (this machine runs a live 1M-context agent session alongside the benchmark). Client-side TTFT figures are contaminated by that traffic; engine per-request timings are unaffected. The surviving observation — one in-flight 200k+ prefill blocks all other clients for up to ~60 s — may interest whoever works on serve scheduling; happy to open a separate issue with logs if useful.

Files: README.md (report), agent_bench.jsonl (20 runs), needles.json (recall), bench_agent_workload.py (script), strata-iq3_s.json (config, paths genericized).

関連リンク

インストール・モデル・リリースへの站内リンク。