Pull requests / #834
#834 bench: community report, RTX 4090 + 32 GB RAM at 512K context (IQ2_XS)
closed · @T-Crypt · 0 评论 · 在 GitHub 查看
BenchmarksNVIDIA / CUDAModels & quantsDocumentation
描述
Community report: **512K context on one RTX 4090 + 32 GB of system RAM** (56 GB total memory), Flash-Next GSQ-RCO IQ2_XS, engine 0.1.39. - `--max-context 524288 --rope-scaling yarn --rope-scale 2 --kv q4_0 --kv-resident 32768 --resident-experts --prefill auto` - A 476,820-token prompt read in **135.5 s (3,411 tok/s)**, decode **122.0 tok/s at depth**, recall correct at all four sizes tested. - On a 32 GB box the RAM floor decides it: with the default resident headroom, 0.1.39 left 1.41 GiB `MemAvailable` at 477K. `STRATA_RESIDENT_HEADROOM_GIB=6` left 3.53 GiB and kept the speed. - `--prefill auto` read long prompts about 3x faster than `--prefill auto:32768` on this box. - Also includes a short single-4090 A/B of #783 (+2 to +5% decode). Limits are in the README: yarn past 262K is experimental, one run per size, recall is a single planted fact rather than needle_bench, and the HF revision of the model files wasn't recorded. Data only: one new directory under `bench/results/`, no engine changes.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。