Pull requests / #1438
#1438 perf(peer): add opt-in byte-capacity-aware expert placement
open · @agorevski · 0 comments · View on GitHub
Setup & installAMD / HIPNVIDIA / CUDADocumentation
Description
## Title <!-- If Applicable, reference the GitHub issue --> Capacity-aware peer expert placement (opt-in) Issue: Not linked to an existing issue. ## Summary <!-- Quick Summary of changes --> Add generalized byte-capacity-aware primary/peer expert placement behind `STRATA_PEER_CAPACITY_AWARE=1`. A large primary cache can swallow the experts moved behind the legacy fixed 8,700-rank boundary; unequal GPUs also need separate byte budgets rather than a common slot-size assumption. The intended benefits are better use of available expert-cache capacity and balancing of hot ranks between owners, not a measured throughput improvement. The profile provides ranks, not measured routing counts: reciprocal square-root rank is only a heat proxy. ## What changed <!-- Specifics on files changed, and what changes were made there --> - `include/strata/core/peer_placement.hpp`: CPU-only planning for heterogeneous, 256-byte-aligned layer blobs, independent byte capacities, peer slot limits, unique ownership, a complete unprofiled tail, and bounded relocation. - `include/strata/core/peer_experts.hpp` and `src/core/peer_experts.cpp`: separate peer scratch preparation, remaining-byte pricing, and expert fill. Opt-in filling excludes actual primary residents, deduplicates candidates, skips oversized blobs, and backfills primary-plan pairs dropped by allocation retries when peer space remains. - `src/program/generate.cpp`: validate opt-in prerequisites, bypass legacy HOT reordering only when enabled, prepare scratch before joint planning, use explicit primary sized-slot ownership, and integrate peer filling. - `tests/core/peer_placement_test.cpp` and `CMakeLists.txt`: register standalone CPU planner coverage for aligned heterogeneous blobs, capacities and slot caps, uniqueness, invalid inputs, cold-tail completion, relocation, and the 21,076-primary-slot / 8,000-profile-rank regression with unequal capacities. - `docs/SECOND_GPU.md`: document opt-in behavior and prerequisites; clarify mapped-host decode transport versus prompt peer-to-peer copies. Without the exact flag value `1` and explicit `--peer-device N`, placement remains unchanged. Opt-in requires a native pack, `--expert-profile`, and a shared, non-elastic expert cache with the CPU pool. It ignores `STRATA_PEER_HOT` / `STRATA_PEER_HOT_AT` and has no GPU-name or architecture-specific selection rule. Existing backend/device requirements still apply. This branch excludes BF16 persistent-prefill-cache changes, their budget reserves, new Gemm/Prefill APIs, and BF16 cache tests; it compiles against the base Gemm/Prefill interfaces. ## Extra Notes <!-- Any extra notes, delete if there are none --> Validation ran from: ```text /home/algore/.copilot/session-state/a7c2bcbb-4264-4eb2-879d-c71a90285b0b/files/pr-worktrees/peer-capacity ``` Exact commands and outcomes: ```bash cmake -S . -B build-peer-capacity-cpu -DSTRATA_ENABLE_CUDA=OFF -DSTRATA_ENABLE_HIP=OFF -DSTRATA_BUILD_TESTS=ON ``` Passed configuration, but selected the environment's Conda compiler. ```bash cmake --build build-peer-capacity-cpu --target peer_placement_test -j2 ``` Failed at linking because Conda's `libstdc++.so` referenced unresolved `dladdr@GLIBC_2.2.5`. The chained CTest command below was consequently not run on that build: ```bash ctest --test-dir build-peer-capacity-cpu -R '^peer_placement_test$' --output-on-failure ``` Retried with a fresh CPU-only system-GCC build: ```bash cmake -S . -B build-peer-capacity-gcc -DCMAKE_C_COMPILER=/usr/bin/gcc -DCMAKE_CXX_COMPILER=/usr/bin/g++ -DSTRATA_ENABLE_CUDA=OFF -DSTRATA_ENABLE_HIP=OFF -DSTRATA_BUILD_TESTS=ON cmake --build build-peer-capacity-gcc --target peer_placement_test -j2 ctest --test-dir build-peer-capacity-gcc -R '^peer_placement_test$' --output-on-failure ``` All three commands passed; CTest passed 1/1 selected test. ```bash /usr/bin/g++ -std=c++20 -O3 -DNDEBUG -Wall -Wextra -DSTRATA_NATIVE_EXPERTS=1 -DSTRATA_PREFILL_FUSED=1 -DSTRATA_PREFILL_MMQ=1 '-DSTRATA_VERSION="0.1.40.3"' -Iinclude -Ithird_party/llama.cpp/ggml/include -isystem /home/algore/miniconda3/include -fsyntax-only src/program/generate.cpp src/core/peer_experts.cpp /usr/bin/g++ -std=c++20 -O3 -DNDEBUG -Wall -Wextra -DSTRATA_NATIVE_EXPERTS=1 -DSTRATA_PREFILL_FUSED=1 -DSTRATA_PREFILL_MMQ=1 '-DSTRATA_VERSION="0.1.40.3"' -Iinclude -Ithird_party/llama.cpp/ggml/include -isystem /home/algore/miniconda3/include -c src/program/generate.cpp -o build-peer-capacity-gcc/generate-host.o /usr/bin/g++ -std=c++20 -O3 -DNDEBUG -Wall -Wextra -DSTRATA_NATIVE_EXPERTS=1 '-DSTRATA_VERSION="0.1.40.3"' -Iinclude -Ithird_party/llama.cpp/ggml/include -isystem /home/algore/miniconda3/include -c src/core/peer_experts.cpp -o build-peer-capacity-gcc/peer-experts-host.o git diff --check ``` All four commands passed. Generate syntax/object compilation reported the pre-existing unused-variable `w` warning. The host checks use CUDA headers but neither link nor execute the CUDA runtime; a full CUDA engine build was not run. No GPU workloads, model servers, or two-GPU runtime tests were run. GPU work is restricted to physical GPU 3, which prevents a peer runtime validation. Actual residency/backfill integration, throughput, and answer parity are therefore **not runtime validated**; only CPU planner coverage and focused host syntax/object compilation are reported. The original checkout, its index/configuration, and the installed engine were not modified.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.