Pull requests / #833

#833 Prefill: batched unbuffered expert reads, close the experts.bin view; opt-in RAM budget with helper caches

closed · @adonizm · 0 评论 · 在 GitHub 查看

BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

描述

## What

Two engine changes, plus one opt-in switch, that speed up prompt reading (prefill) on a PC whose RAM cannot hold all the experts. That is the 32 GB case, where the file tier reads unbuffered.

1. **Batched stager reads** (`src/prefill/prefill.cpp`, `src/core/expert_source.cpp`). Each prompt-path stager thread now claims a run of consecutive jobs. The source reads them with a single `read_direct` batch through the new `ExpertSource::copy_blobs`, which merges neighbouring blobs into requests of up to 32 MiB. Before this, every job was its own ~2 MB request. Set with `STRATA_STAGER_BATCH`: the default is 8 for file-backed blobs and 1 otherwise. `STRATA_STAGER_BATCH=1` restores the old behaviour.
2. **Close the `experts.bin` view once reads are unbuffered** (`FileExpertSource::drop_mapping`, called from `generate.cpp` after the RAM copy is built). While the file is mapped, NTFS serves its unbuffered reads one at a time. After startup the view is only used as a fallback, so it is closed: `mapped_blob` returns nullptr and every file read goes through `read_direct`. `STRATA_KEEP_MAPPING=1` keeps the view open (A/B).
3. **`STRATA_RESIDENT_REMOTE=1`** (experimental, opt-in). Allows `--resident-budget-gib` alongside the CUDA1-3 helper caches (`--expert-cache-device1`). The prompt path runs on CUDA0 and streams from the source, so the RAM budget is what speeds up long prompts, while the helpers serve decode. Without the variable, the existing refusal is unchanged.

## Why

On a 32 GB PC, IQ3_XXS read a 47.6K prompt at 227 tok/s. `STRATA_PREFILL_TIMING=1` showed 86% of the GPU timeline waiting on expert blobs, and drive counters showed ~660 MB/s of mmap page faults. With a RAM budget and unbuffered reads it was still bound by the SSD, at queue depth 1 and ~2.3 GB/s.

A standalone test on the same drive (WD SN580, 2 MiB unbuffered reads at QD32): **2.35 GB/s with experts.bin mapped, 3.57 GB/s without.** With the view closed, the drive runs at an average queue depth of 12 during prefill.

## Measurements

RTX 5070 Ti 16 GB + RTX 3060 12 GB, Ryzen 7 7800X3D, 32 GB RAM, Windows (WDDM), IQ3_XXS (Swift), 47.6K-token prompt, `--prefill auto:32768`. One run per arm unless noted.

| Configuration | Prompt tok/s |
|---|---:|
| Original config (8192 chunks, mmap) | 226.7 |
| `--resident-budget-gib 18`, release 0.1.39 | 1,261 |
| same, this build, `STRATA_STAGER_BATCH=1` | 1,260 |
| batch 4 / 8 / 16 | 1,344 / 1,370 / 1,381 |
| batch 8 + view closed | **1,669** |
| batch 8, view kept (`STRATA_KEEP_MAPPING=1`) | 1,330 |
| `STRATA_RESIDENT_REMOTE=1`, budget 16 GiB, `--pool-workers 7` (two runs) | 1,665 / 1,787 |

Decode, measured on three 600-token answers: the original config gives 33.6 tok/s. A single GPU with the RAM budget gives 35.6 tok/s. The RAM budget with the CUDA1 helper (`STRATA_RESIDENT_REMOTE=1`) gives 43.3 tok/s.

## Notes for review

- `drop_mapping` acts only when `direct_` is non-empty and the source is `experts.bin`. It does nothing for the GGUF-in-place mode, which keeps `maps_` for the token embedding and PLE. `base_` stays non-null only as the "opened" mark. `mapped_blob` checks `unmapped_`, and `close()` skips the unmap when the view is already closed.
- The stager batch is capped at `kRing`, so the earliest unfinished run never waits on a ring buffer it holds itself. A run that mixes in non-transient jobs copies those individually.
- Outputs are not bit-identical to before. A different read order does not change any bytes, but `STRATA_RESIDENT_REMOTE` changes where experts are computed (GPU helper versus CPU), exactly as the helper cache already does. Generated text was checked for coherence, not with a KL comparison.
- Not tested: Linux (`drop_mapping` is Windows-only by design, and `copy_blobs` falls back to the default loop), HIP, and layer splits.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。