贡献 / #1696
#1696 Resident RAM mode on Windows: the start-up fill and the RAM copy read experts.bin unbuffered (ready 42 -> 10 s with 21 GiB free)
open · @sergqwer · 0 评论 · 去 GitHub 看
Setup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux
说明
With `--resident-experts` (what setup writes in the low-RAM mode when the experts the GPU does not hold fit the RAM; the pack then has `experts.bin` from `iq_pack.py --experts-bin`), the start copies two things out of `experts.bin` through its mapped view: - the GPU cache's pre-fill from the expert profile (~19 GB with a 32 GB card); - the RAM copy, `pin_cache_complement` (16.4 GiB here). Every page those copies touch sits in the engine's working set beside the page-locked copy, until #467's trim after the fill and the `VirtualUnlock` after the copy. #467's comment has the result on a real 32 GB PC: 20.7 GiB available before the start, 0.44 GiB after the fill. Windows pages other memory out, and the copies run at page-fault speed (~0.85 GB/s here). **Change** (Windows, `experts.bin` only): - `FileExpertSource::begin_startup_unbuffered` runs once the file tier is chosen, when that tier reads through the file cache. - It closes the view, the section and the buffered handle. NTFS serializes unbuffered reads of a file while it is mapped or open buffered (the reason for `drop_mapping`). - It opens the unbuffered handles (`open_direct`) and reads the first blob both ways. - If the open fails or the two reads differ, the view is mapped again and the start copies through it as before. - The pre-fill and the copy then take the paths #286/#362 already have for an unbuffered tier: - the fill reads `prefetch_pairs` batches of 64, the next one while the current one is copied; - the copy reads `read_direct` batches of 32 blobs on 6 threads; - both merge requests up to 32 MiB. - `end_startup_unbuffered` runs right after the copy. - It maps the view again (`map_view`, with the file size checked). - It closes the unbuffered handles and the stage buffers, and puts the file-tier counters back. - So the file tier (and #577's recheck) is what it was before. - `pin_cache_complement` no longer `VirtualUnlock`s the address range of a closed view. - `STRATA_STARTUP_UNBUFFERED=0` gives the mapped copy back (the A/B arm). - `STRATA_VERIFY_COMPLEMENT=1` prints two checksums: the filled cache slots (read back from VRAM) and the RAM copy (with its offsets). - Unchanged: - Linux (the call is under `_WIN32`); - the GGUF in place, which `begin_startup_unbuffered` refuses: it takes only one mapped `experts.bin`; - a file tier that already reads unbuffered (`--resident-budget-gib` on a small PC); - runs without a RAM copy; - the SYCL port, which has its own copies of these files. **A/B on current main** - Builds: fb58e0db vs this branch, both from the same build script (MSVC, CUDA 13, sm_120). - Machine: RTX 5090, Ryzen 9 9950X3D, 128 GB, Samsung 9100 PRO, Windows 11. - Model: IQ2_XS (ISTA), packed with `iq_pack.py --experts-bin` (33.0 GiB `experts.bin`). - Flags as setup writes them in the low-RAM mode: `--resident-experts --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --max-context 262144 --kv int8`, plus `--vram-reserve-mib 1500 --stop-eos`. - The cache gets ~15,200 slots; the RAM copy is 16.4 GiB. - Prompts: chat (44 tokens in, up to 1,024 out) and the 32K explain prompt (256 out). - Interleaved pairs, order alternating. - "Ready" runs from the process start to the engine's `resident RAM mode:` line, i.e. the copy done. **A 32 GB PC, emulated:** - A RAM ballast (its own process, locked large pages) leaves 21 GiB available: what #467's 32 GB PC had before the start. - 45 s settle after the ballast. - A watchdog stops a run once available RAM stays under 1 GiB for 2 s. - Two series of 3 pairs for each prompt: | | fb58e0db | this PR | paired diff | |---|---|---|---| | ready, chat | 42.5 ± 1.3 s (n4) | 9.6 ± 0.8 s (n6) | **-32.8 ± 1.8 s** (4 pairs) | | ready, 32K | 40.7 ± 0.9 s (n5) | 10.2 ± 0.4 s (n6) | **-30.5 ± 1.3 s** (5 pairs) | | pre-fill (~19 GB) | 26.0 s, 842 MB/s | 3.2 s, 6,914 MB/s | | | RAM copy (16.4 GiB) | 13.7 s | 3.3 s | | | lowest available RAM | 0.25 ± 0.33 GiB | 4.5 ± 1.2 GiB | | | pages written out per run (chat / 32K) | 60,210 / 58,707 | 579 / 3,340 | | | pages read in per run (chat) | 11.7 M | 0.95 M | | | starts the watchdog stopped | 3 of 12 | 0 of 12 | | How many base starts trip the watchdog depends on what the desktop holds at the time. A third series the same hour, with `STRATA_PREFILL_CPU_SHARE=0` in both arms, had the watchdog stop 5 of 6 base starts and none of this PR's. **128 GB, no ballast**, 3 pairs: | | fb58e0db | this PR | |---|---|---| | ready, 32K | 34.6 ± 0.1 s | 9.4 ± 0.2 s | | peak working set | 32.6 ± 0.5 GiB | 18.3 ± 0.2 GiB | | lowest available RAM | 85.5 GiB | 99.9 GiB | | 32K prompt | 5,175 ± 151 ms | 5,186 ± 45 ms (+11 ± 109) | | decode, chat | 10.62 ± 0.37 ms/round | 10.75 ± 0.35 (+0.13 ± 0.21) | | decode, 32K | 11.37 ± 0.10 ms/round | 11.47 ± 0.30 (+0.10 ± 0.20) | **A restart with the file cached**: the worst case for this change, since the mapped copy then reads RAM. Three base starts in a row, then three of this PR: | start | ready | read from the drive | |---|---|---| | fb58e0db, cold | 35.0 s | 33.1 GiB | | fb58e0db, warm (x2) | 13.6 / 13.6 s | 0.2 / 0.1 GiB | | this PR, right after a mapped start | 9.7 s (1.34 s of it purging the cached pages) | 36.9 GiB | | this PR, again (x2) | 8.3 / 8.3 s | 36.9 / 36.9 GiB | **Identity:** - Bit identity uses a budget-bound cache: `--expert-cache 12000 --pcie-frac 0.25`, no ballast. - Tokens are the md5 of the `output :` ids; logits are the md5 of `STRATA_DUMP_FIRST_LOGITS`. | run | tokens | first-token logits | cache checksum | RAM copy checksum | |---|---|---|---|---| | p2k, 200 tokens, fb58e0db (x2) | 3f052d750a | 1e96c370f5 | - | - | | p2k, 200 tokens, this PR (x2) | 3f052d750a | 1e96c370f5 | 4a2f31657ef73b09 (12,595 slots) | 2c1aa305913e5a62 (19.91 GiB) | | p2k, 200 tokens, this PR, `STRATA_STARTUP_UNBUFFERED=0` | 3f052d750a | 1e96c370f5 | 4a2f31657ef73b09 | 2c1aa305913e5a62 | | 32K, 64 tokens, fb58e0db / this PR / `=0` | 47d6bb0797 | 62e9e6e16a | 4a2f31657ef73b09 (PR, `=0`) | 2c1aa305913e5a62 (PR, `=0`) | | 32K, 256 tokens, 24 GiB available (ballast), `--expert-cache 16000`: fb58e0db / this PR / `=0` | 70a7575ec9 | fead174dd7 | 4bf8e7b372672109 (15,438 slots; PR, `=0`) | 81664c8deb808a3b (16.09 GiB; PR, `=0`) | | p2k, the default config (GGUF in place, no RAM copy): fb58e0db / this PR | same | 1e96c370f5 | - | - | `file_expert_source_test` passes. **Risks:** - **The purge.** The first unbuffered open with no buffered handle left purges the file's cached pages. - It took 1.2-2.5 s after a mapped start had left the 33 GiB file cached, 0.3-1.2 s at 21 GiB available, and 3 ms after an unbuffered start. - It is included in the ready times above. - **Every start reads the experts from the drive** (36.9 GiB here), even when the file is cached. - At ~6.9 GB/s that still beats the warm mapped copy (8.3 vs 13.6 s). - On a SATA SSD (~0.5 GB/s) it would take ~70 s where a warm mapped start takes ~14 s. A cold mapped start reads the drive too, in page-fault-sized requests. - Setup writes `--resident-experts` only when the RAM cannot hold the experts with room to spare, so the file is not cached there. - A PC with RAM to spare that passes `--resident-experts` by hand keeps the old path with `STRATA_STARTUP_UNBUFFERED=0`. - Not measured: SATA drives. - **The fallback.** The first blob is read both ways. A failed open or a mismatch maps the view again and copies through it. - After that check there is no mapping to fall back to, as after `drop_mapping` today. - A read that fails later stops the pre-fill with `the profile fill failed at pair N`, as with an unbuffered tier today. In the RAM copy it reports that the copy does not fit, and the soft mode reads those experts from the file. - Through the mapping, the same disk error is an in-page exception. - **The first request right after the start.** #467's trim (`SetProcessWorkingSetSize(-1, -1)` in `pin_cache_complement`) leaves the engine to fault ~1 GiB of its own memory back in after the copy. That happens in both builds: the samples show 60-160 MiB read from the drive every 0.25 s for 2-2.5 s after the `resident RAM mode:` line. - Under the ballast the 44-token chat prompt took 684 ± 356 ms here, against 540 ± 25 ms (n6 / n4); one run took 1.5 s with the CPU share off. - The 32K prompt: +105 ± 326 ms under the ballast. - No difference without the ballast (323 vs 342 ms; 32K +11 ± 109 ms). - The first token still comes ~30 s earlier. - **A layer split.** A stage's cache, and `--expert-cache-per-layer`, fill through `blob_stable` one blob at a time, as with an unbuffered tier today. One card here: not measured. - **Not measured:** - multi-GPU; - HIP on Windows (the same code; not built here); - drives other than this one. - Linux is unchanged by construction: everything this uses outside `_WIN32` already exists there, and the call is not made. The same change in our fork measured, on a 63 GiB NVFP4 pack with 88 GiB available: ready 53.0 -> 17.5 s, peak working set 85 -> 47 GiB, pages written out 1,736 -> 0. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
本站相关内容
相关页面的快捷入口。