Issues / #498

#498 setup: offer the layer split (`--gpus`) for UD-Q4_K_XL when the GGUFs fit in RAM — the engine already runs it without the budget (2x RTX 3090: 31 → 64-78 tok/s)

closed · @mad9home · 1 Kommentare · Auf GitHub

BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsDocumentationLinux

Beschreibung

### Summary

`docs/UNSLOTH_Q4.md` says UD-Q4_K_XL is "one GPU only: with `--gpus` it uses the first one". The reason is the RAM budget mode (`--resident-budget-gib`), which the engine refuses with a layer split (`generate.cpp`: "--resident-cpu-experts does not support layer splits"). Without the budget, the engine already splits UD-Q4_K_XL across two cards, reading the experts that the GPUs do not hold through the OS file cache. On a PC with enough RAM for the GGUFs, this roughly doubles decode speed.

Proposal: when `--gpus` is given (or two usable cards are found) and the RAM can hold the GGUFs plus headroom, setup could drop `--resident-budget-gib` and write `"gpu": [..], "layer_split": "auto"`. That is the mapped mode that #384 already uses for the low-RAM resident variant. It would bridge the gap until the "real resident mode across a split" from #384 lands.

### Setup

- Engine **0.1.35** (tag `v0.1.35`, `d9ab843`), Docker image built from source, sm_86, no vision
- VM: 20 vCPU EPYC 7453 (AVX2), **165 GiB RAM**, 2x RTX 3090 24 GB (PCIe Gen4 x16), driver 595.91.07, Linux
- UD-Q4_K_XL shards from the pinned revision `38bb39e`, SHA-256 verified, mounted read-only

Config: the one that setup wrote for one GPU, edited by hand:

```diff
- "gpu": 0,
+ "gpu": [0, 1],
+ "layer_split": "auto",
  args: ... --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp
            --max-context 262144 --kv int8 --kv-resident 32768
-           --resident-budget-gib 71
```

The engine starts with `layer split across GPUs [0, 1] (auto)`, and the expert cache holds 4,077 experts (11.99 GiB).

### Results

Warm runs (each load gets at least 3 warm-up requests first, because Q4 decode drifts up after a load), 1,024 generated tokens, greedy, thinking off, `strata-bench.py`-style prompts:

| UD-Q4_K_XL | 1x 3090 (setup default, budget 71 GiB) | 2x 3090 (split, no budget) |
|---|---:|---:|
| Decode, 4K prompt | 30.5-32.3 tok/s (median 31.7) | **64-78 tok/s** |
| Decode, 128K prompt (118,745 tok) | ~30 tok/s (engine 0.1.34) | **55.4 tok/s** |
| Prefill, 4K | 843-973 tok/s | 1,155-1,270 tok/s |
| Prefill, 128K | — | 2,192 tok/s (TTFT 54.8 s) |
| Decode expert cache hit rate | ~80 % | 91-94 % |
| VRAM | ~23.8 GB | 23.7 / 23.9 GB |
| Load until `/v1/models` answers | ~2 min | 91 s |

These numbers match PR #417 (2x RTX PRO 4500, UD-Q4_K_XL: 61.5 → 111.5 tok/s short, 70 → 130 tok/s at 128K), which also had to run the split "in the default cache mode (the resident budget cannot be combined with a split)".

### RAM

This is the concern from #384, so I measured `MemAvailable` every 2 s through warm-up, the 4K runs and the 128K run:

- Process RSS ~95 GiB, page cache holds the rest of the 104 GiB of shards
- **Minimum `MemAvailable`: 68 GiB of 165 GiB**, stable, no swapping, no 0-free episode

So on a PC with RAM well above the GGUF size, the mapped mode is safe. The 0-free case in #384 was 47 GB of RAM with a ~70 GB model. A threshold of roughly "GGUF size + 24 GB" (the same headroom the budget already uses) would separate the two cases.

### Possible setup change

1. If UD-Q4_K_XL is chosen with `--gpus` (or `offer_together` finds a pair) **and** RAM ≥ GGUF size + headroom: write the split config without `--resident-budget-gib` (mapped mode), and print a note that the experts come through the file cache.
2. Otherwise: keep today's behaviour (first card, budget), with the note.
3. In `UNSLOTH_Q4.md`, change "One GPU only" to describe the condition.

Happy to test a branch on this box (2x 3090, 165 GiB), or to open a PR for the setup part if you prefer.

Mehr auf der Site

Links zu Install, Modellen, Releases.