Pull requests / #945
#945 setup: the low-RAM and KV-streaming decisions as functions
closed · @pipeob0 · 0 comments · View on GitHub
Setup & installNVIDIA / CUDAModels & quants
Description
### Why A front-end that wants to tell the user *what would change* on an installed model needs the same answer setup.py gives when it installs: is this model in the low-RAM mode here, does it stream its KV cache at this context. Step 5 and step 7 decided both inline, so anyone else had to re-derive the RAM math and would drift. ### What - `low_ram_wanted(model, ram, choice)` and `kv_streaming_wanted(model, ctx, kv, ram, choice)` (plus `kv_streaming_ram_gb`), named next to `low_ram_fits`; step 5 and step 7 now call them. The streaming branches keep their messages and say they are `kv_streaming_wanted`. - Behaviour unchanged: `low_ram_wanted` is the old `budget is None and (...)`, since `budget` is set only for a size that has a RAM budget. ### Tests `test_setup_risk.RulesAgreeWithTheWrittenConfig`: 54 mocked installs (3 sizes x 2 contexts x 3 `--kv` x 3 `--kv-streaming`) and 24 more (2 PCs x 4 `--low-ram`), each comparing the written config's `--kv-resident` / `--resident-experts` / `--mmap-experts` with the function. `python -m unittest tools.test_setup_risk` - no GPU, no download. ### AI Written by an agent (pi) running on Strata's own local model (`qwen3.8-flash-next-iq3_s`), reviewed and tested by hand on an RTX 5060 Ti 16 GB / Ryzen 7 5700X3D / 64 GB PC with engine 0.1.39.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.