Pull requests / #44
#44 generate: probe the real host->device bandwidth for pcie_frac
closed · merged 2026-09-28 · @pipeob0 · 0 comentários · No GitHub
Setup & installNVIDIA / CUDAModels & quants
Descrição
**Independent of #43 — no stacking needed.** It applies on plain `main`: it adds one function and replaces the `pcie_frac` default line, while #43 touches only the `expert kernels run on` notice a few lines below. Verified by cherry-picking this commit onto `origin/main` with no conflicts. ### Why `pcie_frac`'s defaults (0.55 native / 0.2 canonical) were measured on a PCIe 4.0 **x16** link (~26 GB/s). The comment right above them says the copy engine's share has to shrink with the link — but nothing measured the link, and a x8 card in a x8 slot is very common: it is what a 16 GB card with a big expert cache lands on on most AM4/AM5 boards. There the share was being set to 2× what the wire can carry, so the GPU waits for DMA that has not arrived. The only user-facing knob was `--pcie-frac`, i.e. a per-user guess. ### What Probe the host→device bandwidth **once at startup, the way the engine actually uses it** — DMA reads from pinned host memory, 4 × 256 MiB, copy engine warmed — then scale the default by it: ```c o.pcie_frac = std::min(base, std::max(0.05, base * (bw / 26.0))); ``` Clamped so it can only shrink the share relative to the measured link and never below 0.05, and never above the current default (a x16 machine keeps exactly today's behaviour). `--pcie-frac` still wins when given, and if the probe cannot run at all the default is kept and the log says so. On this box (RTX 5060 Ti 16 GB at PCIe 4.0 x8): ``` strata generate: PCIe probe: 14.1 GB/s host->device -> pcie_frac 0.30 (default 0.55) ``` ### Cost and risk One-time at startup: ~1 GB of copies over a 14 GB/s link ≈ 75 ms, plus a 256 MiB pinned buffer and a 256 MiB device allocation that are freed again before the arena is built. If either allocation fails, the probe returns "could not run" and nothing changes. The probe runs after CUDA context setup, so it measures the same path the expert fetches use. Behaviour change is confined to machines whose link is slower than x16 PCIe 4.0 — the population the current default is wrong for. It is also self-documenting: the log now states the measured bandwidth and both the probed and the stock value, so a report can say whether the share was probed or guessed. Not claimed: an end-to-end speed number for this. It moves the copy-engine share on x8 links; the round-time effect on this machine was inside the noise of the other work in progress, so I am not attaching a number that the method (±5% acceptance-driven round counts) cannot support.
No site
Links install, modelos, releases.