Pull requests / #1038

#1038 decode: STRATA_PCIE_BALANCE=1 picks each layer's PCIe count from measured costs (opt-in)

closed · @sergqwer · 0 commentaires · Sur GitHub

BenchmarksServer & APIAMD / HIPNVIDIA / CUDADocumentation

Description

The verify window reads `pcie_frac` of a layer's missed experts over PCIe. The link probe sets that share, so the CPU's speed never enters it. On a slow CPU the GPU spins on the pool's rows while the link sits mostly idle.

This adds an opt-in, `STRATA_PCIE_BALANCE=1`. Off (the default), nothing changes: no new kernel is captured and the share is the probe's.

### What it does

- **Stamps.** `wait_flag_ge_stamped` is the same flag spin kernel, which also writes the GPU clock at three points: as the layer's plan arrives (flag A), as the PCIe part starts (flag B) and as the GPU starts waiting for the CPU's rows. This adds no extra launch, and the stamped kernels are captured only with the switch on.
- **Fits.** The host already times its pool per step. After each window, least squares with forgetting (~200 layers) fit two lines:
  - the pool: `a + c·n_cpu + d·n_pcie`;
  - the GPU: `g0 + g·n_pcie`.

  `d` is there because the pool and the link read the same RAM. On a DRAM-bound pool, a PCIe copy slows the pool by about what the expert it took off the CPU saved, so moving misses buys nothing there. A plain per-expert balance got this wrong and was ~4% slower on such a box.
- **Choice.** A layer reads over PCIe the `m` of its misses that minimizes `max(GPU, CPU)`. Every 64th layer that would keep all its misses on one side moves one to the other, so neither fit goes stale. Without that, a cold first sample once sent everything over PCIe for good.
- **Scope.** One GPU, no batch slots and no `--pcie-frac`. A serve request's own `pcie_frac` pauses the switch for that request. The decode summary prints the fitted costs.

### Measured

IQ2_XS (ISTA's GSQ-RCO), RTX 5090 + Ryzen 9 9950X3D, chat prompt decoded to the end of the turn (`--stop-eos`), automatic cache:

| | off (probe 0.55) | `STRATA_PCIE_BALANCE=1` |
|---|---|---|
| one pool worker on the AVX2 kernels (`STRATA_FORCE_AVX2=1 --pool-workers 1 --no-host-worker`) | 16.77 / 17.17 ms a round, 154 / 149 tok/s | **12.98 / 12.82 ms, 198 / 200 tok/s** |
| the 9950X3D's 16 workers, 3 pairs | 12.09 / 12.03 / 12.45 | 12.02 / 11.99 / 12.50 |

- **One worker.** The fit reads the CPU at 42-60 us an expert and the GPU at 28 us a PCIe expert, with `d` 26-28 us, so 0.8 of the misses move.
- **16 workers.** The pool is DRAM-bound (`d` 0.5-0.6 of `c`) and few misses move.

Off, the logits and the 32 greedy tokens equal v0.1.39 on IQ2_XS: 2K prompt, fixed cache, identical bytes.

The slow-CPU case is emulated on this machine by fewer workers and the AVX2 kernels, not run on such a CPU. A 6-core or a Zen 2 box (the README's Ryzen 9 3900X) sits somewhere between those two rows. That is the case the fit is for: it finds the box's own split instead of the probe's. The HIP builds take the clock `gpu_stamp` already uses (`wall_clock64`), but I have no AMD card to run them on.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

Sur le site

Liens install, modèles, releases.