Issues / #746

#746 Regarding 128GB Memory Optimization Strategies

closed · @Susan985 · 5 comentarios · En GitHub

Benchmarks

Descripción

My hardware configuration includes 128GB of 8-channel DDR4 3200MHz memory. When running the UD-Q4_K_XL model, the system frequently reads data from the SSD, even though up to 60GB of memory remains available. Is there a way to keep the model resident in memory to achieve optimal performance? Although removing the `--resident-budget-gib` parameter did not completely resolve the issue, performance improved significantly: prefill speeds exceeded 1900 tokens/s, and decode speeds reached 65–100 tokens/s. My goal is to have the model run entirely in memory without relying on SSD disk mapping.

En el sitio

Enlaces a install, modelos, releases.