Pull requests / #377
#377 HIP: time the PCIe probe on the host on Windows
closed · @BlueKingMuch · 0 comentários · No GitHub
AMD / HIPNVIDIA / CUDAWindowsLinux
Descrição
The PCIe probe sets the share of the missed experts the GPU reads over the link (`--pcie-frac`): the native default 0.55 from 20 GB/s up, less on a slower link (#44). On Windows HIP it timed its copies with events that do not bracket them there, so it read impossible speeds and every link kept the x16 share. ## What changes - On Windows HIP builds, `probe_pcie_h2d_gbps` times the same 4 x 256 MiB copies on the host clock, between two `cudaDeviceSynchronize` calls; the warm-up copy is synchronized before the clock starts. The 1 GiB takes about 40 ms on this card, so the synchronize around it hardly matters. - CUDA and Linux HIP builds keep the events (`#if defined(STRATA_USE_HIP) && defined(_WIN32)`). - The layer split's per-GPU probe calls the same function, so it gets the same timing. ## Measured RX 6800 on PCIe 4.0 x16, Windows 11, HIP SDK 7.2: | | events (before) | host clock (this PR) | |---|---|---| | probe reading | 3,348 to 25,999 GB/s over 24 engine starts (median about 19,000) | 27.2 GB/s in all 4 starts | | resulting `pcie_frac` | 0.55 | 0.55 | A standalone copy of the probe on that card shows why: its events measure 0.13 ms for the 1 GiB, the host clock 43 ms (25.1 GB/s). The same test reads 27.2 GB/s for a copy from registered memory and 27.0 GB/s for a kernel reading mapped pinned memory, so 27.2 GB/s is what this link carries. In one engine start with `GPU_FLUSH_ON_EXECUTION=1` (the runtime submits every command right away) the events read 26.9 GB/s. On this card the share is 0.55 either way, so nothing changes in speed here. On a slower link the share now shrinks as #44 intended, e.g. to 0.28 at 13 GB/s (about what PCIe 3.0 x16 carries); not tested on such a link. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_015Ld5YDzuKW1hgMVwwgRZDP
No site
Links install, modelos, releases.