Issues / #737
#737 [enhancement] #498 follow-up: UD-Q4_K_XL split gate — allow confirm/override just below 135 GB total RAM
closed · @biteric2000 · 2 comentários · No GitHub
BenchmarksMulti-GPUNVIDIA / CUDAModels & quants
Descrição
**TL;DR:** the multi-GPU split for `UD-Q4_K_XL` is hard-blocked when total RAM is under ~135 GB, but its real peak is ~97 GB. On a 128 GB box it runs fine (split, ~100 tok/s, no errors). Could this become a confirm/override instead of a hard no?
Thanks for Strata — running a 125B MoE on a desktop is a great bit of work. Small ask, no urgency.
## The gate
```python
def unsloth_split_need_gb(model="UD-Q4_K_XL") -> float:
return MODELS[model]["download_gb"] + UNSLOTH_RAM_LEFT_GB # ~111.3 + 24 = ~135 GB
```
- `ram_gb()` is **total** RAM, so a 128 GB box fails `128 < 135.3` regardless of usage.
- The #498 note says it's safe at 165 GiB with MemAvailable ≥ 68 GiB (peak ~97 GB) — so 135 GB is full file size + 24 GB margin, a deliberate worst case.
Conservative by design, which is fine. Only gap: a 128 GB box with a small OS footprint can't reach split at all.
## It works at 128 GB (measured)
Bypass = 3 config-JSON changes: `gpu: [0,1]`, `layer_split: "auto"`, and removing `--resident-budget-gib`.
2× RTX 4090 (24 GB), i9-14900K (AVX2), 128 GB DDR5, Win11, NVMe. Engine 0.1.37, `--max-context 262144`, `--kv int8`.
- RAM ~89 GB used / ~39 GB free (no #384 0-free)
- decode ~90–115 tok/s · expert cache hit 90.6% → 97.9% · no OOM / `cudaHostRegister` errors
#498 measured 64–78 tok/s on 2× 3090 / 165 GiB; I get ~100 tok/s on 2× 4090 / 128 GiB. One data point — anecdotal.
## Small idea (any one, or none)
1. Make `ram < need` a `confirm_risk` prompt ("Go on anyway?") — matches the "recommends, never forces" note.
2. An optional flag, e.g. `--unsloth-ram-left-gib N`.
3. Or just document the manual path for sub-135 GB boxes.
Happy to leave it as-is if the current behaviour is the right call.
No site
Links install, modelos, releases.