Pull requests / #1548

#1548 decode: STRATA_PCIE_BALANCE=1 picks each layer's PCIe count from measured costs (refile of #1038 by @sergqwer, on 0.1.40.3)

open · @stuchapin909 · 0 评论 · 在 GitHub 查看

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsSecurityWindows

描述

## Summary
A refile of #1038 by @sergqwer, which GitHub closed when `main`'s history was rewritten (not a rejection; your note there asked for a new branch from the new `main`). The commit is theirs, cherry-picked onto 0.1.40.3 with `-x`, authorship kept. Opt-in: `STRATA_PCIE_BALANCE=1` picks each layer's PCIe count from measured CPU and GPU costs; unset, nothing changes. The method, the scope (one GPU, no batch slots, no `--pcie-frac`) and the original measurements are in #1038.

## What changed
Nothing beyond #1038's commit, except one conflict: in `include/strata/kernels/verify_kernels.hpp` the comment above `gpu_stamp` changed on `main` (the PDL sentence) while #1038 added `wait_flag_ge_stamped` above it. Both are kept: the new declaration, then `main`'s current `gpu_stamp` comment. `wait_flag_ge_stamped` launches like `wait_flag_ge` (no PDL), as on `main`.

## Measured
The case #1038 did not cover: a strong CPU on a slow link, where the probe's share is too high. i9-9920X (Skylake-X, 12 cores, AVX-2 expert kernels), 96 GB DDR4, Windows 11, one RTX 3090 on PCIe 3.0 x16 (probe 12.3 GB/s -> `pcie_frac 0.34`), built with `STRATA_PORTABLE=ON`. IQ3_S (GSQ-RCO), expert cache 2,454 slots (4.71 GiB, `--vram-reserve-mib 12300`, about a 12 GB card), `--spec 4 --spec-min-p 0.5 --kv k8v4 --kv-resident 32768 --max-context 262144`, one prompt ("Explain how to compute 17*19 + 23*29 by hand, then give the integer."), greedy, 256 tokens, `strata generate --tokens-file`, a fresh process per run.

| | switch unset | `STRATA_PCIE_BALANCE=1` |
| --- | --- | --- |
| 2026-10-07, #1038's head cherry-picked (alternating with `--pcie-frac 0.25`) | 52.83, 51.90 tok/s | 57.04, 57.41 tok/s |
| 2026-10-08, this commit | 52.51 tok/s | 54.82, 55.75 tok/s |
| experts per layer read over PCIe | 5.55-5.57 | 4.45-4.78 |

+9.2% and +5.3% over the probe share in the two sittings. The fitted model printed `CPU 162-173 us + 57-65 an expert + 0 a PCIe one, GPU 145-178 us + 146-150 a PCIe expert`. The best fixed share on this PC, by hand, is 0.25 (58.0-58.2 tok/s, 4.1 experts a layer over PCIe), so the switch gets most of the way there without tuning. The same tokens in every run.

On this PC's two-card layer split the share makes no measurable difference (automatic vs 0.25, 3 starts each), in line with the switch's one-GPU scope.

## Extra Notes
@sergqwer: the change is yours; if you would rather refile it yourself, I will close this one.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。