Pull requests / #1533

#1533 Freing up RAM on machines with more VRAM, by not holding VRAM experts in RAM. (And swapping from VRAM to RAM when need other expert on GPU)

open · draft · @gfifgfifofich · 0 Kommentare · Auf GitHub

BenchmarksServer & APINVIDIA / CUDAModels & quants

Beschreibung

### What this is

A suggestion from a box that could not run Unsloth's UD-IQ4_XS quant of Qwen3.8-Flash-Next until these were fixed.
Three commits, all in resident RAM mode with a layer split:

| commit | one line |
|---|---|
| `fab4985` | `--vram-reserve-mib` took `atoi`, so a non-number became 0 and the engine died with `prefill gemm: cublasCreate: cuBLAS status 1` and no mention of the flag |
| `97320f5` | `--adapt-swaps` above 96 did nothing: `min(o.adapt_swaps, 96)` in three places, including the exchange-buffer reserve |
| `a1b0cd0` | the prompt loan moves experts instead of copying them, and the complement build releases only its own pages |

Opening it as a draft because the third commit asks a design question I can't settle from one machine.

### The machine

Two RTX 5070 Ti (16 GiB each), 62.7 GiB of system RAM, model files on ntfs3 on external NVMe (which is why the
drive numbers below look like 1.5-2 GB/s and not NVMe numbers). Model:
`Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf`, 24576 experts, 56.3 GiB of expert bytes, largest blob 3.48 MB.

### 1. The complement build left every page it read in the file cache

The release after copying a layer ran only for `experts.bin` (`role_ptr_.empty()`) — the comment says *"the GGUF in
place leaves its pages to the OS"*. With the GGUF mapped in place, which is the usual layout, the build reads the
complement (39.2 GiB here) into the pinned arena and keeps every page it read in the file cache as well: a 62.7 GiB
machine holds ~78 GiB of expert bytes halfway through the load and swaps before the copy finishes. This is the
difference between the box booting and not.

Adding the release for the mapped case exposed a second problem. It dropped each layer's whole gate/up/down range,
which includes the experts a GPU cache holds — and those are exactly the bytes the prompt path reads back every
chunk, so the refill went from a page-cache hit to a drive read. It now walks the layer's experts and releases only
the runs the arena owns (`offsets[i] != kNoComplement`).

5.5k-token prompt, same model, same box, same config:

```
release missing:              33.3 s    338 tok/s prompt processing   24.14 GiB read
release added (whole layer):   7.5 s    946 tok/s                     14.48 GiB
release narrowed to the copy:  ~6 s    ~1.7k tok/s
```

**The ask:** the files are on ntfs3, where `posix_fadvise` returns success (the new warning never fired) but I
cannot see whether it really drops the pages. On ext4/btrfs this should be a much cleaner measurement, and I'd
rather this number came from someone else's box than from mine.

### 2. The prompt loan is a transfer, not a copy

The RAM copy is the complement of the GPU caches by construction. The prompt path borrows GPU cache slots as
prefill buffers, so the experts sitting in those slots have no copy in RAM, and the loan has to be refilled from
the GGUF. Measured on this model: **3317 blobs, 9.36 GiB of drive reads per prompt cycle** — CUDA0's 2073 slots
plus CUDA1's 2074 of 3468. A phase-separated probe says the same from the other side: 1.32 GiB/s of drive reads
during a prompt, **0.00 GiB/s during decode**. All of the drive traffic is the loan.

`pin_cache_complement` now reserves that space and leaves it empty (`lend_offsets_`, `lend_reserved_`); `lend_save`
moves an expert into its reserved space as the slot goes out, and `lend_release` clears it once the refill lands.
Both are virtual no-ops on `ExpertSource` and are called from the plain generate path and the serve path (where
`refill_wait` releases only that participant's pairs). Holding the space at all is opt-in (`--lend-keeps-ram`),
because reserving RAM for experts that sit in a VRAM slot is the duplication that put this box into swap. Related:
#403 turned the loan off whenever a budget was given; the ceiling should decide, so a budget now caps complement
and reserve together, and a later stage's loan gets reserved space too.

**Stated plainly, because it is the honest part:** on this box the reserve engaged for **298 of 1657 slots**
(`lend save: 298 experts moved ... (1359 had none and stay on the files)`) — the free-RAM ceiling took the room
before the reserve got any. The prompt speed in section 1 is not from this commit. This is the mechanism that lets
a machine with more RAM stop reading the drive during a prompt; here it is mostly not engaged.

### 3. The design question

The rule the code now enforces is: *an expert is in RAM exactly when no GPU slot holds it* — with the loan as the
one exception, and it has to be asked for. That is the opposite of what a mirror would do (keep a RAM copy of the
cached experts too, so a prompt never reads the drive), and I removed that from my local tree because on this box
`--resident-budget-gib 50` plus a 4 GiB mirror pinned 54 GiB on a 62.7 GiB machine, page-lock registration came
back partial, and the engine produced **plausible tokens from the wrong expert with nothing in the log**. I would
like that failure mode to be impossible by construction rather than by a flag default, and I think this diff does
that. If you disagree, the argument is worth having before I rework it.

### What was tested, and where

Built clean on `main` (`6674a00`, 0.1.40.4), CUDA, sm_120, Release.

```
file_expert_source_test      PASS
conversation_cache_test      4198 checks passed
conversation_memory_test     23 checks passed
resident_plan_parity         OK
```

`ctest` cannot be believed on this box right now: both cards are held at 15.6/16.3 GiB by a running engine, so
`cudaMalloc: out of memory`. 36 of 64 fail on this branch and 34 of 64 on unmodified `main` at the same moment;
the 7 that differ (`elementwise_parity`, `kv_hybrid_parity`, `kv_steps_parity`, `resident_plan_parity`,
`rope_parity`, `sampler_parity`, `verify_parity`) all pass when run on their own. So: **the CUDA suite has not
been run in a clean state for this diff** and I'd rather say that than imply it passed.

The end-to-end numbers above were measured on `v0.1.40.1` with the same diff, not on 0.1.40.4 — the port across
those 267 commits touched `expert_source.cpp` in three places (the role-offset arithmetic, since `role_fd_`/
`role_off_` are mine and not upstream, and an `OnDevice` include) and it compiles and passes the tests listed
above, but I have not run a 5.5k prompt through a 0.1.40.4 engine with this in it.

### Not in this PR

A VRAM→host read-back tier (serve cache-held experts over PCIe instead of off the drive) is in my local tree, off
by default and self-verifying against the file bytes. It is not here because it is a separate question and its
failure mode — plausible bytes for the wrong expert, nothing in the log — deserves its own argument.

Mehr auf der Site

Links zu Install, Modellen, Releases.