Issues / #746

#746 Regarding 128GB Memory Optimization Strategies

closed · @Susan985 · 5 评论 · 在 GitHub 查看

Benchmarks

描述

My hardware configuration includes 128GB of 8-channel DDR4 3200MHz memory. When running the UD-Q4_K_XL model, the system frequently reads data from the SSD, even though up to 60GB of memory remains available. Is there a way to keep the model resident in memory to achieve optimal performance? Although removing the `--resident-budget-gib` parameter did not completely resolve the issue, performance improved significantly: prefill speeds exceeded 1900 tokens/s, and decode speeds reached 65–100 tokens/s. My goal is to have the model run entirely in memory without relying on SSD disk mapping.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。