Pull requests / #100

#100 fix: honor native per-layer layouts in file-backed expert loading

closed · @midhatn · 0 コメント · GitHub で見る

NVIDIA / CUDAWindows

本文

Native expert packs can have different blob sizes in different layers. `FileExpertSource` still sizes `experts.bin` and addresses every blob using the canonical fixed stride, so `--mmap-experts` rejects these native packs or uses the wrong layout.

This uses the already-loaded native layout's per-layer offsets, strides and total size. It validates geometry and extents, retains the canonical path, and unmaps the actual mapping length on POSIX. The pinned-arena path is unchanged.

Validation on Windows/MSVC, against current main (`c1e9033`, v0.1.20):

- Built the engine and `native_mmap_test`.
- The regression fails against the unmodified loader at `native mapping opens`, and passes with this patch.
- Covers mixed native layer strides, every byte of each fixture blob, axis bounds, geometry mismatch, truncated files, close, and canonical compatibility.

This is a loader correctness fix, not a throughput claim. The POSIX path has not been run locally. It is related to #80's memory goal, but does not add tiering or change its policy; keeping native file addressing correct should be useful independently.

Reproduce with a CUDA-enabled build: build target `native_mmap_test`, then run the resulting executable. It uses synthetic files and requires no model download.

関連リンク

インストール・モデル・リリースへの站内リンク。