Pull requests / #44

#44 generate: probe the real host->device bandwidth for pcie_frac

closed · merged 2026-09-28 · @pipeob0 · 0 comentários · No GitHub

Setup & installNVIDIA / CUDAModels & quants

Descrição

**Independent of #43 — no stacking needed.** It applies on plain `main`: it adds one function and replaces the `pcie_frac` default line, while #43 touches only the `expert kernels run on` notice a few lines below. Verified by cherry-picking this commit onto `origin/main` with no conflicts.

### Why

`pcie_frac`'s defaults (0.55 native / 0.2 canonical) were measured on a PCIe 4.0 **x16** link (~26 GB/s).
The comment right above them says the copy engine's share has to shrink with the link — but nothing measured
the link, and a x8 card in a x8 slot is very common: it is what a 16 GB card with a big expert cache lands on
on most AM4/AM5 boards. There the share was being set to 2× what the wire can carry, so the GPU waits for DMA
that has not arrived. The only user-facing knob was `--pcie-frac`, i.e. a per-user guess.

### What

Probe the host→device bandwidth **once at startup, the way the engine actually uses it** — DMA reads from
pinned host memory, 4 × 256 MiB, copy engine warmed — then scale the default by it:

```c
o.pcie_frac = std::min(base, std::max(0.05, base * (bw / 26.0)));
```

Clamped so it can only shrink the share relative to the measured link and never below 0.05, and never above
the current default (a x16 machine keeps exactly today's behaviour). `--pcie-frac` still wins when given, and
if the probe cannot run at all the default is kept and the log says so.

On this box (RTX 5060 Ti 16 GB at PCIe 4.0 x8):

```
strata generate: PCIe probe: 14.1 GB/s host->device -> pcie_frac 0.30 (default 0.55)
```

### Cost and risk

One-time at startup: ~1 GB of copies over a 14 GB/s link ≈ 75 ms, plus a 256 MiB pinned buffer and a 256 MiB
device allocation that are freed again before the arena is built. If either allocation fails, the probe
returns "could not run" and nothing changes. The probe runs after CUDA context setup, so it measures the same
path the expert fetches use.

Behaviour change is confined to machines whose link is slower than x16 PCIe 4.0 — the population the current
default is wrong for. It is also self-documenting: the log now states the measured bandwidth and both the
probed and the stock value, so a report can say whether the share was probed or guessed.

Not claimed: an end-to-end speed number for this. It moves the copy-engine share on x8 links; the round-time
effect on this machine was inside the noise of the other work in progress, so I am not attaching a number
that the method (±5% acceptance-driven round counts) cannot support.

No site

Links install, modelos, releases.