Pull requests / #1696

#1696 Resident RAM mode on Windows: the start-up fill and the RAM copy read experts.bin unbuffered (ready 42 -> 10 s with 21 GiB free)

open · @sergqwer · 0 commentaires · Sur GitHub

Setup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

Description

With `--resident-experts` (what setup writes in the low-RAM mode when the experts the GPU does not hold fit the RAM; the pack then has `experts.bin` from `iq_pack.py --experts-bin`), the start copies two things out of `experts.bin` through its mapped view:
- the GPU cache's pre-fill from the expert profile (~19 GB with a 32 GB card);
- the RAM copy, `pin_cache_complement` (16.4 GiB here).

Every page those copies touch sits in the engine's working set beside the page-locked copy, until #467's trim after the fill and the `VirtualUnlock` after the copy. #467's comment has the result on a real 32 GB PC: 20.7 GiB available before the start, 0.44 GiB after the fill. Windows pages other memory out, and the copies run at page-fault speed (~0.85 GB/s here).

**Change** (Windows, `experts.bin` only):
- `FileExpertSource::begin_startup_unbuffered` runs once the file tier is chosen, when that tier reads through the file cache.
  - It closes the view, the section and the buffered handle. NTFS serializes unbuffered reads of a file while it is mapped or open buffered (the reason for `drop_mapping`).
  - It opens the unbuffered handles (`open_direct`) and reads the first blob both ways.
  - If the open fails or the two reads differ, the view is mapped again and the start copies through it as before.
- The pre-fill and the copy then take the paths #286/#362 already have for an unbuffered tier:
  - the fill reads `prefetch_pairs` batches of 64, the next one while the current one is copied;
  - the copy reads `read_direct` batches of 32 blobs on 6 threads;
  - both merge requests up to 32 MiB.
- `end_startup_unbuffered` runs right after the copy.
  - It maps the view again (`map_view`, with the file size checked).
  - It closes the unbuffered handles and the stage buffers, and puts the file-tier counters back.
  - So the file tier (and #577's recheck) is what it was before.
- `pin_cache_complement` no longer `VirtualUnlock`s the address range of a closed view.
- `STRATA_STARTUP_UNBUFFERED=0` gives the mapped copy back (the A/B arm).
- `STRATA_VERIFY_COMPLEMENT=1` prints two checksums: the filled cache slots (read back from VRAM) and the RAM copy (with its offsets).
- Unchanged:
  - Linux (the call is under `_WIN32`);
  - the GGUF in place, which `begin_startup_unbuffered` refuses: it takes only one mapped `experts.bin`;
  - a file tier that already reads unbuffered (`--resident-budget-gib` on a small PC);
  - runs without a RAM copy;
  - the SYCL port, which has its own copies of these files.

**A/B on current main**
- Builds: fb58e0db vs this branch, both from the same build script (MSVC, CUDA 13, sm_120).
- Machine: RTX 5090, Ryzen 9 9950X3D, 128 GB, Samsung 9100 PRO, Windows 11.
- Model: IQ2_XS (ISTA), packed with `iq_pack.py --experts-bin` (33.0 GiB `experts.bin`).
- Flags as setup writes them in the low-RAM mode: `--resident-experts --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --max-context 262144 --kv int8`, plus `--vram-reserve-mib 1500 --stop-eos`.
  - The cache gets ~15,200 slots; the RAM copy is 16.4 GiB.
- Prompts: chat (44 tokens in, up to 1,024 out) and the 32K explain prompt (256 out).
- Interleaved pairs, order alternating.
- "Ready" runs from the process start to the engine's `resident RAM mode:` line, i.e. the copy done.

**A 32 GB PC, emulated:**
- A RAM ballast (its own process, locked large pages) leaves 21 GiB available: what #467's 32 GB PC had before the start.
- 45 s settle after the ballast.
- A watchdog stops a run once available RAM stays under 1 GiB for 2 s.
- Two series of 3 pairs for each prompt:

| | fb58e0db | this PR | paired diff |
|---|---|---|---|
| ready, chat | 42.5 ± 1.3 s (n4) | 9.6 ± 0.8 s (n6) | **-32.8 ± 1.8 s** (4 pairs) |
| ready, 32K | 40.7 ± 0.9 s (n5) | 10.2 ± 0.4 s (n6) | **-30.5 ± 1.3 s** (5 pairs) |
| pre-fill (~19 GB) | 26.0 s, 842 MB/s | 3.2 s, 6,914 MB/s | |
| RAM copy (16.4 GiB) | 13.7 s | 3.3 s | |
| lowest available RAM | 0.25 ± 0.33 GiB | 4.5 ± 1.2 GiB | |
| pages written out per run (chat / 32K) | 60,210 / 58,707 | 579 / 3,340 | |
| pages read in per run (chat) | 11.7 M | 0.95 M | |
| starts the watchdog stopped | 3 of 12 | 0 of 12 | |

How many base starts trip the watchdog depends on what the desktop holds at the time. A third series the same hour, with `STRATA_PREFILL_CPU_SHARE=0` in both arms, had the watchdog stop 5 of 6 base starts and none of this PR's.

**128 GB, no ballast**, 3 pairs:

| | fb58e0db | this PR |
|---|---|---|
| ready, 32K | 34.6 ± 0.1 s | 9.4 ± 0.2 s |
| peak working set | 32.6 ± 0.5 GiB | 18.3 ± 0.2 GiB |
| lowest available RAM | 85.5 GiB | 99.9 GiB |
| 32K prompt | 5,175 ± 151 ms | 5,186 ± 45 ms (+11 ± 109) |
| decode, chat | 10.62 ± 0.37 ms/round | 10.75 ± 0.35 (+0.13 ± 0.21) |
| decode, 32K | 11.37 ± 0.10 ms/round | 11.47 ± 0.30 (+0.10 ± 0.20) |

**A restart with the file cached**: the worst case for this change, since the mapped copy then reads RAM. Three base starts in a row, then three of this PR:

| start | ready | read from the drive |
|---|---|---|
| fb58e0db, cold | 35.0 s | 33.1 GiB |
| fb58e0db, warm (x2) | 13.6 / 13.6 s | 0.2 / 0.1 GiB |
| this PR, right after a mapped start | 9.7 s (1.34 s of it purging the cached pages) | 36.9 GiB |
| this PR, again (x2) | 8.3 / 8.3 s | 36.9 / 36.9 GiB |

**Identity:**
- Bit identity uses a budget-bound cache: `--expert-cache 12000 --pcie-frac 0.25`, no ballast.
- Tokens are the md5 of the `output  :` ids; logits are the md5 of `STRATA_DUMP_FIRST_LOGITS`.

| run | tokens | first-token logits | cache checksum | RAM copy checksum |
|---|---|---|---|---|
| p2k, 200 tokens, fb58e0db (x2) | 3f052d750a | 1e96c370f5 | - | - |
| p2k, 200 tokens, this PR (x2) | 3f052d750a | 1e96c370f5 | 4a2f31657ef73b09 (12,595 slots) | 2c1aa305913e5a62 (19.91 GiB) |
| p2k, 200 tokens, this PR, `STRATA_STARTUP_UNBUFFERED=0` | 3f052d750a | 1e96c370f5 | 4a2f31657ef73b09 | 2c1aa305913e5a62 |
| 32K, 64 tokens, fb58e0db / this PR / `=0` | 47d6bb0797 | 62e9e6e16a | 4a2f31657ef73b09 (PR, `=0`) | 2c1aa305913e5a62 (PR, `=0`) |
| 32K, 256 tokens, 24 GiB available (ballast), `--expert-cache 16000`: fb58e0db / this PR / `=0` | 70a7575ec9 | fead174dd7 | 4bf8e7b372672109 (15,438 slots; PR, `=0`) | 81664c8deb808a3b (16.09 GiB; PR, `=0`) |
| p2k, the default config (GGUF in place, no RAM copy): fb58e0db / this PR | same | 1e96c370f5 | - | - |

`file_expert_source_test` passes.

**Risks:**
- **The purge.** The first unbuffered open with no buffered handle left purges the file's cached pages.
  - It took 1.2-2.5 s after a mapped start had left the 33 GiB file cached, 0.3-1.2 s at 21 GiB available, and 3 ms after an unbuffered start.
  - It is included in the ready times above.
- **Every start reads the experts from the drive** (36.9 GiB here), even when the file is cached.
  - At ~6.9 GB/s that still beats the warm mapped copy (8.3 vs 13.6 s).
  - On a SATA SSD (~0.5 GB/s) it would take ~70 s where a warm mapped start takes ~14 s. A cold mapped start reads the drive too, in page-fault-sized requests.
  - Setup writes `--resident-experts` only when the RAM cannot hold the experts with room to spare, so the file is not cached there.
  - A PC with RAM to spare that passes `--resident-experts` by hand keeps the old path with `STRATA_STARTUP_UNBUFFERED=0`.
  - Not measured: SATA drives.
- **The fallback.** The first blob is read both ways. A failed open or a mismatch maps the view again and copies through it.
  - After that check there is no mapping to fall back to, as after `drop_mapping` today.
  - A read that fails later stops the pre-fill with `the profile fill failed at pair N`, as with an unbuffered tier today. In the RAM copy it reports that the copy does not fit, and the soft mode reads those experts from the file.
  - Through the mapping, the same disk error is an in-page exception.
- **The first request right after the start.** #467's trim (`SetProcessWorkingSetSize(-1, -1)` in `pin_cache_complement`) leaves the engine to fault ~1 GiB of its own memory back in after the copy. That happens in both builds: the samples show 60-160 MiB read from the drive every 0.25 s for 2-2.5 s after the `resident RAM mode:` line.
  - Under the ballast the 44-token chat prompt took 684 ± 356 ms here, against 540 ± 25 ms (n6 / n4); one run took 1.5 s with the CPU share off.
  - The 32K prompt: +105 ± 326 ms under the ballast.
  - No difference without the ballast (323 vs 342 ms; 32K +11 ± 109 ms).
  - The first token still comes ~30 s earlier.
- **A layer split.** A stage's cache, and `--expert-cache-per-layer`, fill through `blob_stable` one blob at a time, as with an unbuffered tier today. One card here: not measured.
- **Not measured:**
  - multi-GPU;
  - HIP on Windows (the same code; not built here);
  - drives other than this one.
  - Linux is unchanged by construction: everything this uses outside `_WIN32` already exists there, and the call is not made.

The same change in our fork measured, on a 63 GiB NVFP4 pack with 88 GiB available: ready 53.0 -> 17.5 s, peak working set 85 -> 47 GiB, pages written out 1,736 -> 0.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

Sur le site

Liens install, modèles, releases.