Pull requests / #129
#129 core: add optional shared expert arena backing
closed · @rhgo1749 · 0 comentários · No GitHub
Multi-GPUNVIDIA / CUDAModels & quantsWindowsLinux
Descrição
Follow-up to #127, limited to the shared-arena piece you invited as a PR. ## What this changes - Adds an optional `--shared-expert-arena FILE` backing for the resident expert arena. - On Linux, `PinnedArena` uses a MAP_SHARED file with a 4 KiB header when that option is set. - The header records the arena size and a pack fingerprint; an existing backing from a different pack is refused before it can be used. - The existing anonymous/hugetlb allocation path is unchanged when the option is absent. - Existing CUDA host registration, sliced registration fallback, resident locking, expert loading, hot cache, and numerical paths are unchanged. - The CLI notes that the backing belongs on `/dev/shm`, not ordinary SSD storage. The fork used an mmap interception wrapper to prove the idea. This version makes the backing explicit in `PinnedArena`/`ArenaExpertSource` instead, so unrelated mappings are not affected. ## Validation - Rebased onto exact `v0.1.26` (`ac8b251`). - `pinned_shared_test`: same-pack mappings observe the same bytes; different-pack and size mismatches are refused. - Full `strata` CUDA build passes locally with CUDA 13.4 / sm_120. - `--mmap-experts` + `--shared-expert-arena` is rejected with exit 2. - Real 3 x RTX 5070 Ti / IQ3_S / 262144-context-per-lane validation passed: all three processes mapped the same shared backing and three concurrent generation requests completed successfully. - The broader multi-process/multi-GPU serving measurements and architecture motivation are in #127. ## Deliberate non-goals - No request router or multi-GPU supervisor in this PR. - No leader/follower population protocol yet; callers still coordinate startup/population. - The backing path must be dedicated to one compatible expert arena; callers also own its lifetime and cleanup. - Linux MAP_SHARED only for now. A Windows named-file-mapping backend can implement the same backing contract separately. This is intended to be the smallest upstreamable primitive needed for the request-per-GPU design from #127.
No site
Links install, modelos, releases.