Pull requests / #1564
#1564 setup: UD-Q4_K_XL / UD-IQ4_XS keep their RAM budget on a layer split (#642)
open · @ShevchenkoVadim · 0 コメント · GitHub で見る
BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
本文
## Title Issue: related to #1145 (the comment there: setup is stale for UD-IQ4_XS on several GPUs, a hand-built config was needed) ## Summary Since engine 0.1.40 (fe9c10ca, #848) the resident RAM copy works on a layer split - every stage's cache is left out of it and the rest is ranked by the whole expert profile - and `--resident-budget-gib` runs through the same code (`pin_cache_complement`), so the engine no longer refuses the budget on a split. `split_mmap()` got the matching setup exemption in 43c8cfa9, but the RAM budget's setup paths still assume the old engine (#498): a start of a UD-Q4_K_XL / UD-IQ4_XS config on several GPUs deletes `--resident-budget-gib` and asks for ~135 GB of RAM, and `--gpus` at setup drops the budget too. This keeps the budget on a split from `RESIDENT_SPLIT_ENGINE`, gated exactly like #642. **The recommendations do not change** (golden configs unchanged): one GPU stays the default; the split is used only when asked for. ## What changed - `setup.py` - `split_budget()`: returns without touching the config when `resident_split()` (as `split_mmap()` does). - `offer_together()`: a budget config is offered both cards at any RAM, with a message that the budget is kept; the default answer stays "n". - `unsloth_together()`: with `--gpus` (also with an explicit `--resident-budget-gib`) or "2" when asked, the cards go together with the budget; asked as before with one GPU recommended; `--yes` alone keeps one GPU. - install: `q4_split` (the "split without the budget" path) only before `RESIDENT_SPLIT_ENGINE`, so `--resident-budget-gib` stays in the args. - `tools/test_setup_unsloth.py`: the #498 tests run with `MIN_ENGINE` pinned to 0.1.39 (as `StartOnSeveralGpus` does for #642); three new tests cover the kept budget at install (`--gpus`, explicit budget, the "2" answer) and at a start (`--gpus` at 165 and 64 GB, the offer). - `docs/UNSLOTH_Q4.md`, `docs/MULTI_GPU.md`: the budget on a split. ## Extra Notes - Tests: `python -m unittest discover -s tools -p "test_setup_*.py"` on Linux (the strata Docker image, `STRATA_EXECV` unset): 421 tests; the 2 failures in `test_setup_engine_hash` are the same on main without this change (no ready-made 0.1.41 engine to download from the test box). - Measured on real hardware, engine 0.1.41, this branch's setup.py starting the config: RTX 3090 (PCIe x16) + RTX 3080 Ti (x4), i9-14900K, 91 GB RAM, UD-Q4_K_XL, budget 64 GiB, 262K context, `--no-prefill-borrow --prefill 4096`, images on: | | 3090 alone | 3090 + 3080 Ti, budget kept | |---|---|---| | layers / GPU cache | 0-47 / 3,886 slots | 0-1 on the x4 card (678 slots), 2-47 on the 3090 (4,182) | | experts in RAM | 60.4 GiB | 57.6 GiB, no file-tier reads per request | | decode, 2,000 tokens | 38.4 tok/s | 45.0 / 45.1 tok/s (two runs) | | 20K prompt | ~18 s | ~18 s | Answers checked by hand (arithmetic, a list, an image). Before this change the same config on two cards lost its budget at every start. - #1145 ran `--resident-budget-gib 18` on a 3-GPU split with engine 0.1.40 for 2 hours without crashes, NaN windows or stalls - the engine side of this. - Not changed: the default stays one GPU for these models. On this box the split was +17% decode; whether two cards should become the recommendation (as #642 did for the low-RAM mode) is your call. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
関連リンク
インストール・モデル・リリースへの站内リンク。