Pull requests / #89
#89 loader: bulk-read the expert arena and the MTP drafter (MSVC splits ifstream reads into 4095-byte calls)
closed · merged 2026-09-29 · @dannychirkov · 0 コメント · GitHub で見る
NVIDIA / CUDAModels & quantsWindowsLinux
本文
# loader: bulk-read the expert arena and the MTP drafter
## What this fixes
On a Windows/MSVC build, `std::ifstream::read` is not one read. `basic_filebuf::xsgetn`
(__msvc_filebuf.hpp) loops over
```cpp
constexpr size_t _Read_size = 4095;
while (_Read_size < _Count_s) { fread(_Ptr, 1, 4095, _Myfile); }
```
so every streaming read is chopped into 4095-byte `fread` calls. For the IQ3_XXS pack
(one 42.9 GB `experts.bin`, arena ~40 GiB) that is ~11.4 million reads, and the load is
dominated by per-call overhead, not by the disk. Measured on this machine: **4095.8 bytes
per operation** over ~11.4M operations, 0.03 GiB/s effective read rate, 1222 s for the
expert phase. The same defect was present in the drafter path (`src/core/mtp.cpp`:
`read_file` of `dense.bin` and the 512 x 64 MiB expert loop).
`src/core/expert_source.cpp` has the same `std::ifstream` read in its GGUF branch
(`load_experts_gguf`); it is not touched here because the pack path is the one this
report measures, but it is the same pattern and probably wants the same fix.
## What changed
* `src/core/pinned.cu`, `include/strata/core/pinned.hpp`, `src/core/expert_source.cpp`,
`include/strata/core/load_main.cpp` - read the arena with `fopen`/`fread` and an explicit
64-bit seek (`_fseeki64`), with a short-read check that fails the load (`LoadStats.ok`,
`LoadStats.error`) instead of silently continuing with partial data.
* `src/core/mtp.cpp` - the same for the drafter: the whole-file `dense.bin` read and the
512-expert loop (`_fseeki64`/`_ftelli64`).
The bytes read, the call sites and the behaviour on Linux/glibc are unchanged; only the
read path is.
## Measurements
One machine, one config (64K context, `--expert-cache 2500`, `--kv int8`, MTP on), cold
start, only the binary differs between the three arms. Boundaries are the *same* for
every arm and every number below:
* **READY** - the engine's own `READY at N s`, from process start (the harness reads the
same instant as `ready_s_engine_clock`).
* **READY (launcher)** - wall clock from the launcher process to the server answering.
* **expert phase** - the engine's `expert load phases:` line (the original has no such
timer; its number is derived from its own `loaded 39.97 GiB at 0.03 GiB/s`, i.e.
39.97 / 0.0327).
* **bytes/op** - the harness's telemetry watchdog, sampled inside the expert phase.
| arm | READY | READY (launcher) | expert phase | expert rate | bytes/op |
|---|---|---|---|---|---|
| 0.1.15 original | 1248.66 s | 1249 s | 1222 s (derived) | 0.03 GiB/s | 4096 |
| + arena fix | 209.55 s | 217.2 s | 47.29 s | 0.85 GiB/s | 8388608 |
| + drafter fix | 72.36 s | 80.2 s | 46.64 s | 0.86 GiB/s | 8388608 |
(The expert phase reads 39.97 GiB; the rate in the engine's own line is 39.97 GiB / phase
wall, while the harness also reports the *summed* per-thread read time, 217.86 s over 6
threads = 78% of the phase wall, i.e. 1.10 GiB/s while actually reading.)
**This says nothing about generation speed.** Decode was not measured in these arms and
this change does not touch it; the only generation run here is a short prompt to prove the
engine serves requests.
## Verification
* A differential stand compiles the *old* and the *new* body of the read function side by
side, with the bodies taken verbatim out of the source: 8/8 offsets reproduce the
original bytes, including offsets above 4 GiB (layer 0 @ 0, layer 6 @ 4 928 307 200,
layer 47 @ 41 798 860 800, a 3 MiB + 615 B block), 48/48 layers match, and the
reassembled files are byte-identical to the originals (external `sha256sum` of the
reassembled files matches the files themselves).
* A short read is refused: reading past EOF returns `ok = 0` with
`short read in layer 0: got 6815744 of 6819840 B at offset 42906157056`.
* The drafter stand: dense.bin (110.7 MiB) 3.40 s -> 0.07 s, experts.bin (675 MiB)
1.43 s -> 0.47 s, a truncated file (700 000 000 of 707 788 800 B) and a missing file are
both refused, and the reassembled bytes match the files' `sha256sum`.
* The real engine was then run on the pack, same config, only the binary changed.
* These two commits were rebased onto `main` (v0.1.20) and **compile clean** there with the same
toolchain (MSVC 14.44.35207 / CUDA 13.0 / Ninja, Release), so the patch is not tied to the
0.1.15 tree the numbers above come from. The numbers themselves were taken on 0.1.15+these
two patches; only the binary differed between arms.
## Caveats
* The two passes over the pack showed the same speed for the original binary, which is
consistent with a limited file cache but does not by itself prove that the read was
fully cold. The fix was not derived from that observation; it was found in the STL
source and then confirmed by the differential stand above.
* This PR does not include the local phase timers used to take the numbers above (they
were local instrumentation, and `src/program/generate.cpp` has changed upstream since),
nor a version-stamp change in `CMakeLists.txt`.
## Environment
Windows 11, MSVC 14.44.35207 / cl 19.44.35229 (Ninja generator), CUDA 13.0 V13.0.48,
CMake 4.3.1, Ryzen 9 3900X, 64 GB DDR4-3200, PCIe 3.0 x8/x8, RTX 5060 Ti 16 GB
(sm_120) + GTX 1050 Ti, pack: ISTA-DASLab GSQ-RCO IQ3_XXS (2 shards + PLE, 3.06 bit/w,
48 layers, 512 experts).
関連リンク
インストール・モデル・リリースへの站内リンク。