Pull requests / #1134

#1134 bench: community report — RTX 5090, IQ3_S at 1M context (YaRN), agent-style workload (re-submission of #466)

closed · @gravitomagnetic · 0 评论 · 在 GitHub 查看

BenchmarksServer & APINVIDIA / CUDAModels & quantsDocumentation

描述

Re-submitting #466 onto the cleaned history — that PR was auto-closed by the force-push per @Niko1221's note, and its branch was built on the old history so it could not be reopened. This branch starts from the new `main` and carries the same content plus everything that lived in the closed PR's comments, folded into the README:

- **Original 0.1.31 report** (bench table, warm-reuse table, Note A on checkpoint tax, needle recall 5/5 up to 1,048,265 tokens via `tools/needle_bench.py`).
- **0.1.39 re-probe**: the prefill regression documented in the report is closed — 4K 2,786 / 32K 3,589 / 128K 3,533 t/s, beating the 0.1.31 golden baseline at 32K (+9%) and 128K (+5%).
- **Preemption A/B** (the head-of-line scenario from @j-luwierski's analysis, measured): 125.2K-token cold prefill colliding with a 49–92-token ping, both request flavors mirroring real provider profiles. FIFO B-wait 49.2 s (thinking) / 31.5 s (no-think) vs `--batch 2` 9.2 s / 4.8 s — **5.3× / 6.5×**, with BYIELD interleaving visible in the serve log and the cost to A quantified (−2% prefill, −13% decode, +2 s wall).
- **VRAM budget finding for 32 GB cards at 1M context**: `--batch 4` OOMs at load; `--batch 2` needs `--expert-cache 6500 --vram-reserve-mib 995` (at 7500/700 it loads but dies on the first request needing slot 2). Suggested as a line in BATCHING.md.
- **Observability note**: batched requests return `prompt_n: 0` / null client-side timings while the serve log has segment truth.
- Raw per-run JSON for both the bench and the four preemption legs included.

No code changes — data + docs only, same as the original. Thanks for the porting offer in the closing note, but a clean re-PR is simpler now that we have the updated content.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。