Issues / #1145

#1145 Resident RAM mode on a layer split — 3 GPUs, UD-IQ4_XS, 40 GB RAM (engine 0.1.40)

open · @chiangww · 2 commentaires · Sur GitHub

BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsWindowsLinux

Description

**Resident RAM mode on a layer split — 3 GPUs, UD-IQ4_XS, 40 GB RAM (engine 0.1.40)**

Answering the "30-minute soak of the resident RAM mode on a split (#848)" call. Note this is 3 GPUs, one of them over USB4 — beyond the 2-GPU case you measured.

**Rig**
- 7700X running Ubuntu 24.04.5 LTS (GNU/Linux 7.0.14-15-pve x86_64)
- RTX 5070 Ti (PCIe5 x8, probe 28.9 GB/s) + RTX 5060 Ti (PCIe4 x8, 14.4 GB/s) + RTX 5060 Ti on USB4 (3.8 GB/s) = 48GB VRAM
- LXC container: 40 GB RAM, models on local NVMe
- UD-IQ4_XS, `--layer-split 30,40`, `--kv k8v4`, `--max-context 204800`, `--resident-experts`, conversation cache 12 GiB / 16 slots
- Workload: 3 agent conversations, 110–120K tokens each, alternating turns

**A/B: full complement (31.58 GiB) vs `--resident-budget-gib 18`**

| | complement 31.58 GiB | budget 18 GiB |
|---|---|---|
| fresh prompt (0 reused) | 436 tok/s (119,730 tok) | **1,369 tok/s** (112,970 tok) |
| reused-tail read | 121 tok/s (9,432 tok) | 359 tok/s (2,776 tok) |
| decode at ~120K | 56–61 tok/s | 35–44 tok/s |
| decode hit rate | 89.5–91.7 % | 82.2–88.3 % |
| PCIe / other-GPU share | 6.0–7.3 % | 6.4–9.4 % |
| conversation cache | refused: "skip parking (physical RAM admission; need … plus 2560 MiB floor)" | parks 113,185 tok in 860 ms, 2.24 GB, 0 evictions |
| complement copy at startup | 67 s cold | 3 s warm |

Reading: with the full complement, prompt streaming misses the pinned complement and reads the GGUF at NVMe speed (436 tok/s × 3.48 MB/blob ≈ 4.5 GB/s), and the 31.58 GiB page-lock starves the conversation cache. Budget 18 GiB (8,056 experts by profile rank) covers the prompt hot set: prefill 3.1× faster and parking works again, at ~30 % decode (adaptive swaps evict outside the complement — `files … blobs` grows ~30 k reads per request). For reference, IQ3_S on this rig in plain mmap mode (0.1.39) did 1,268 tok/s at 126K — UD-IQ4_XS with budget 18 matches it despite 31 % larger blobs.

**Two observations**
1. `7438 verified additional-GPU experts remain on the mmap fallback` — the CUDA1/2 cache contents are not in the resident complement, so adaptive swaps involving them read files. Should the complement (or the budget) cover them?
2. `the file tier reads unbuffered … the file cache cannot keep them` — with budget 18 there are 18.5 GiB free, yet the whole 37.4 GiB is read O_DIRECT. Would a budgeted run gain from letting the page cache hold the overflow just past the budget, instead of the file?

⚠️ No crashes, NaN windows or stalls in 2h of interactive use so far; soak continues.

**setup.py: the #498 guard is stale for 0.1.40.** `split_budget()` strips `--resident-budget-gib` on any multi-GPU start and demands download_gb + 24 ≈ 118 GB for UD-IQ4_XS, pointing 40 GB rigs at "start it on one GPU". But with 0.1.40 (#848) the budget runs on a layer split — verified above, this whole test is that config. The strip path seems to need the same engine-version gate `split_mmap()` already uses (RESIDENT_SPLIT_ENGINE). My config was hand-built from the IQ3_S one because setup would not offer it; it works, and the budget is what makes 40 GB viable.

Sur le site

Liens install, modèles, releases.