Pull requests / #864

#864 perf: opt-in resident expert exchange buffer rotation

closed · @CC-David-CC · 0 コメント · GitHub で見る

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows

本文

Resident RAM exchanges copy each evicted expert from a temporary buffer into the promoted expert's old slot. This patch rotates buffer ownership after the existing reader/transfer completion boundary, removing that host copy.

Clean extraction of `aa6e0e3` + `1a50d91` and their tests, as requested in #689, based on upstream `6f32ec0`. Includes the focused guide and one Q2_0 A/B. Q8 PLE support and the broader performance experiments remain separate.

`STRATA_EXCHANGE_ROTATE=1` opts in; default off. It requires equal-size expert blocks and fully pinned/mapped resident and exchange buffers. Other layouts keep the existing copy path. Presets, placement policy and GPU synchronization are unchanged; configurations without the resident RAM complement do not use rotation.

### Validation

- Clean code at `18a30ad775bce86e99a54e6dc552ccb72c66d16e`; same binary in both arms. RTX PRO 6000 Blackwell 96GB, Ryzen 7950X, GCC 13.3, CUDA 13.2.
- Standard GSQ-RCO Q2_0 and matching IQ4_NL PLE; both shards SHA-256 verified. Same 1,024 input tokens, fresh engine per arm, **all 1,024 generated token IDs identical**. FP16 KV, greedy target-only, zero drafts and prompt reuse.
- CPU ownership checks (12,304 exchanges), both CTest checks, and CUDA copy/rotation/pageable-fallback fixtures passed. ASan/UBSan passed; Compute Sanitizer reported 0 errors.

| Rotation | Decode tokens/s | Resident exchanges | Avoided host-copy payload | Output IDs |
|---|---:|---:|---:|---|
| `0` | 104.90 | 3,802 | 0 bytes | 1,024, identical |
| `1` | 113.44 | 3,802 | 5,255,884,800 bytes | 1,024, identical |

The GPU cache was deliberately capped at 12,000 slots to exercise the 16.19 GiB RAM complement; Q2 can fit fully on this card. One pair, rates exclude startup/prefill, EOS sentinel forces the exact length. No general speedup, model-quality or speculative-path parity claim. AMD and Windows GPU execution remain untested.

[Guide and verification commands](https://github.com/CC-David-CC/Strata-a5500/blob/perf/exchange-rotation-only/docs/EXCHANGE_ROTATION.md) | [Receipt](https://github.com/CC-David-CC/Strata-a5500/blob/perf/exchange-rotation-only/bench/results/exchange-rotation-ab.json) | [Raw token IDs](https://github.com/CC-David-CC/Strata-a5500/blob/perf/exchange-rotation-only/bench/results/exchange-rotation-token-ids.json)

関連リンク

インストール・モデル・リリースへの站内リンク。