Issues / #746
#746 Regarding 128GB Memory Optimization Strategies
closed · @Susan985 · 5 评论 · 在 GitHub 查看
描述
My hardware configuration includes 128GB of 8-channel DDR4 3200MHz memory. When running the UD-Q4_K_XL model, the system frequently reads data from the SSD, even though up to 60GB of memory remains available. Is there a way to keep the model resident in memory to achieve optimal performance? Although removing the `--resident-budget-gib` parameter did not completely resolve the issue, performance improved significantly: prefill speeds exceeded 1900 tokens/s, and decode speeds reached 65–100 tokens/s. My goal is to have the model run entirely in memory without relying on SSD disk mapping.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。