Pull requests / #1438

#1438 perf(peer): add opt-in byte-capacity-aware expert placement

open · @agorevski · 0 commentaires · Sur GitHub

Setup & installAMD / HIPNVIDIA / CUDADocumentation

Description

## Title
<!-- If Applicable, reference the GitHub issue -->
Capacity-aware peer expert placement (opt-in)

Issue: Not linked to an existing issue.

## Summary
<!-- Quick Summary of changes -->
Add generalized byte-capacity-aware primary/peer expert placement behind
`STRATA_PEER_CAPACITY_AWARE=1`. A large primary cache can swallow the experts
moved behind the legacy fixed 8,700-rank boundary; unequal GPUs also need
separate byte budgets rather than a common slot-size assumption.

The intended benefits are better use of available expert-cache capacity and
balancing of hot ranks between owners, not a measured throughput improvement.
The profile provides ranks, not measured routing counts: reciprocal
square-root rank is only a heat proxy.

## What changed
<!-- Specifics on files changed, and what changes were made there -->
- `include/strata/core/peer_placement.hpp`: CPU-only planning for heterogeneous,
  256-byte-aligned layer blobs, independent byte capacities, peer slot limits,
  unique ownership, a complete unprofiled tail, and bounded relocation.
- `include/strata/core/peer_experts.hpp` and `src/core/peer_experts.cpp`: separate
  peer scratch preparation, remaining-byte pricing, and expert fill. Opt-in
  filling excludes actual primary residents, deduplicates candidates, skips
  oversized blobs, and backfills primary-plan pairs dropped by allocation
  retries when peer space remains.
- `src/program/generate.cpp`: validate opt-in prerequisites, bypass legacy HOT
  reordering only when enabled, prepare scratch before joint planning, use
  explicit primary sized-slot ownership, and integrate peer filling.
- `tests/core/peer_placement_test.cpp` and `CMakeLists.txt`: register standalone
  CPU planner coverage for aligned heterogeneous blobs, capacities and slot
  caps, uniqueness, invalid inputs, cold-tail completion, relocation, and the
  21,076-primary-slot / 8,000-profile-rank regression with unequal capacities.
- `docs/SECOND_GPU.md`: document opt-in behavior and prerequisites; clarify
  mapped-host decode transport versus prompt peer-to-peer copies.

Without the exact flag value `1` and explicit `--peer-device N`, placement
remains unchanged. Opt-in requires a native pack, `--expert-profile`, and a
shared, non-elastic expert cache with the CPU pool. It ignores
`STRATA_PEER_HOT` / `STRATA_PEER_HOT_AT` and has no GPU-name or
architecture-specific selection rule. Existing backend/device requirements
still apply.

This branch excludes BF16 persistent-prefill-cache changes, their budget
reserves, new Gemm/Prefill APIs, and BF16 cache tests; it compiles against the
base Gemm/Prefill interfaces.

## Extra Notes
<!-- Any extra notes, delete if there are none -->
Validation ran from:

```text
/home/algore/.copilot/session-state/a7c2bcbb-4264-4eb2-879d-c71a90285b0b/files/pr-worktrees/peer-capacity
```

Exact commands and outcomes:

```bash
cmake -S . -B build-peer-capacity-cpu -DSTRATA_ENABLE_CUDA=OFF -DSTRATA_ENABLE_HIP=OFF -DSTRATA_BUILD_TESTS=ON
```

Passed configuration, but selected the environment's Conda compiler.

```bash
cmake --build build-peer-capacity-cpu --target peer_placement_test -j2
```

Failed at linking because Conda's `libstdc++.so` referenced unresolved
`dladdr@GLIBC_2.2.5`. The chained CTest command below was consequently not run
on that build:

```bash
ctest --test-dir build-peer-capacity-cpu -R '^peer_placement_test$' --output-on-failure
```

Retried with a fresh CPU-only system-GCC build:

```bash
cmake -S . -B build-peer-capacity-gcc -DCMAKE_C_COMPILER=/usr/bin/gcc -DCMAKE_CXX_COMPILER=/usr/bin/g++ -DSTRATA_ENABLE_CUDA=OFF -DSTRATA_ENABLE_HIP=OFF -DSTRATA_BUILD_TESTS=ON
cmake --build build-peer-capacity-gcc --target peer_placement_test -j2
ctest --test-dir build-peer-capacity-gcc -R '^peer_placement_test$' --output-on-failure
```

All three commands passed; CTest passed 1/1 selected test.

```bash
/usr/bin/g++ -std=c++20 -O3 -DNDEBUG -Wall -Wextra -DSTRATA_NATIVE_EXPERTS=1 -DSTRATA_PREFILL_FUSED=1 -DSTRATA_PREFILL_MMQ=1 '-DSTRATA_VERSION="0.1.40.3"' -Iinclude -Ithird_party/llama.cpp/ggml/include -isystem /home/algore/miniconda3/include -fsyntax-only src/program/generate.cpp src/core/peer_experts.cpp
/usr/bin/g++ -std=c++20 -O3 -DNDEBUG -Wall -Wextra -DSTRATA_NATIVE_EXPERTS=1 -DSTRATA_PREFILL_FUSED=1 -DSTRATA_PREFILL_MMQ=1 '-DSTRATA_VERSION="0.1.40.3"' -Iinclude -Ithird_party/llama.cpp/ggml/include -isystem /home/algore/miniconda3/include -c src/program/generate.cpp -o build-peer-capacity-gcc/generate-host.o
/usr/bin/g++ -std=c++20 -O3 -DNDEBUG -Wall -Wextra -DSTRATA_NATIVE_EXPERTS=1 '-DSTRATA_VERSION="0.1.40.3"' -Iinclude -Ithird_party/llama.cpp/ggml/include -isystem /home/algore/miniconda3/include -c src/core/peer_experts.cpp -o build-peer-capacity-gcc/peer-experts-host.o
git diff --check
```

All four commands passed. Generate syntax/object compilation reported the
pre-existing unused-variable `w` warning. The host checks use CUDA headers
but neither link nor execute the CUDA runtime; a full CUDA engine build was
not run.

No GPU workloads, model servers, or two-GPU runtime tests were run. GPU work
is restricted to physical GPU 3, which prevents a peer runtime validation.
Actual residency/backfill integration, throughput, and answer parity are
therefore **not runtime validated**; only CPU planner coverage and focused
host syntax/object compilation are reported. The original checkout, its
index/configuration, and the installed engine were not modified.

Sur le site

Liens install, modèles, releases.