Pull requests / #134
#134 setup: count the 3-bit models' 262K context against RAM, not a fixed 90 GB
closed · @architectds · 0 comments · View on GitHub
Setup & installModels & quants
Description
The second piece from #103, as you suggested. `setup.py` cuts IQ3_S and IQ3_XXS asked for 262K down to 128K on any machine under 90 GB of RAM. With KV streaming, the context costs mostly RAM, so this PR counts the need instead: the experts' arena, plus the context's 8-bit KV (~13.7 KB a token, the same figure setup uses for `--kv-resident`), plus 24 GB of room for everything else. | RAM | IQ3_XXS at 262K (needs 70.5 GB) | IQ3_S at 262K (needs 77.9 GB) | |---|---|---| | 64 GB PC (68.7 GB) | 128K, as before | 128K, as before | | 83.5 GiB box (89.6 GB) | 262K (was cut to 128K) | 262K (was cut to 128K) | Measured: a Colab A100-40G (83.5 GiB of RAM) runs IQ3_S at 262K with images in 56 GiB. The check still only applies above 128K, and it still uses the 8-bit KV figure, because the KV precision is asked after the context. Nothing else changes.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.