Pull requests / #1325

#1325 file tier (Windows): read each expert from experts.bin and a mirror on a second drive at once

open · @malloc32 · 0 Kommentare · Auf GitHub

BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsWindows

Beschreibung

**AI-written.** This PR was prepared by an AI (Claude Opus 5.5, via Claude Code) at the submitter's request; the code and the measurements (including the read-latency micro-benchmark) were done by the AI on the submitter's machine. The idea of using both drives came from the submitter, who reviewed it before opening.

**Depends on #1323** (the #833 port is this branch's first commit); review that one first.

**Update 2026-10-07:** rebased onto 0.1.40.2 (`e8ca9af`) together with #1323; no conflicts in this branch's own commits. The mirror has been in daily use here since (checked again after a reboot).

## What

- `STRATA_EXPERTS_MIRROR=<path>`: a byte-identical copy of `experts.bin` on another drive. Each unbuffered request is read as two halves, one per drive, in flight together. Before use, the copy's size and 16 sampled 64 KiB blocks (first and last included) are compared with the original; a stale copy is refused and the reads stay on one drive.
- `STRATA_DIRECT_SPLIT_KIB=K`: a request larger than K KiB is read as K-KiB slices issued together (alternating drives with a mirror).

Windows only, `experts.bin` only. Without the variables nothing changes.

## Measured

Details in `bench/results/2026-10-07-windows-experts-mirror/`.

- One ~2 MB expert, unbuffered: 482 us from one drive (Samsung 990 PRO); 288 us with its two halves read from the 990 PRO and a KIOXIA EXCERIA PRO at once.
- With a RAM budget and unbuffered reads (IQ3_S, 2x RTX 5060 Ti): decode **53.7 +- 0.8 -> 58.4 tok/s** (+8.7%) with mirror + 512 KiB slices; prompts +3-5%.

Cost: a second copy of `experts.bin` (50 GB here), which has to be copied again whenever the pack changes.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Mehr auf der Site

Links zu Install, Modellen, Releases.