Pull requests / #129

#129 core: add optional shared expert arena backing

closed · @rhgo1749 · 0 コメント · GitHub で見る

Multi-GPUNVIDIA / CUDAModels & quantsWindowsLinux

本文

Follow-up to #127, limited to the shared-arena piece you invited as a PR.

## What this changes

- Adds an optional `--shared-expert-arena FILE` backing for the resident expert arena.
- On Linux, `PinnedArena` uses a MAP_SHARED file with a 4 KiB header when that option is set.
- The header records the arena size and a pack fingerprint; an existing backing from a different pack is refused before it can be used.
- The existing anonymous/hugetlb allocation path is unchanged when the option is absent.
- Existing CUDA host registration, sliced registration fallback, resident locking, expert loading, hot cache, and numerical paths are unchanged.
- The CLI notes that the backing belongs on `/dev/shm`, not ordinary SSD storage.

The fork used an mmap interception wrapper to prove the idea. This version makes the backing explicit in `PinnedArena`/`ArenaExpertSource` instead, so unrelated mappings are not affected.

## Validation

- Rebased onto exact `v0.1.26` (`ac8b251`).
- `pinned_shared_test`: same-pack mappings observe the same bytes; different-pack and size mismatches are refused.
- Full `strata` CUDA build passes locally with CUDA 13.4 / sm_120.
- `--mmap-experts` + `--shared-expert-arena` is rejected with exit 2.
- Real 3 x RTX 5070 Ti / IQ3_S / 262144-context-per-lane validation passed: all three processes mapped the same shared backing and three concurrent generation requests completed successfully.
- The broader multi-process/multi-GPU serving measurements and architecture motivation are in #127.

## Deliberate non-goals

- No request router or multi-GPU supervisor in this PR.
- No leader/follower population protocol yet; callers still coordinate startup/population.
- The backing path must be dedicated to one compatible expert arena; callers also own its lifetime and cleanup.
- Linux MAP_SHARED only for now. A Windows named-file-mapping backend can implement the same backing contract separately.

This is intended to be the smallest upstreamable primitive needed for the request-per-GPU design from #127.

関連リンク

インストール・モデル・リリースへの站内リンク。