Pull requests / #1713
#1713 sycl/Arc: stepped verify window (opt-in STRATA_VERIFY_STEPPED) - CPU experts served between graph segments
open · @aslater3 · 0 commentaires · Sur GitHub
BenchmarksSetup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsLinux
Description
Stacked on #1602 (Alchemist setup + staging-scratch fix). The first new commit here is the last one. ## What it does `STRATA_VERIFY_STEPPED=1` is opt-in and off by default, so other GPUs are unaffected. It changes how the doorbell verify window serves CPU experts on Arc A-series. Today the window is one graph. GPU kernels spin on host flags while the CPU pool serves the off-GPU experts inside it. On an A770 that spinning costs about 100–170 ms of GPU wait per window. With `STRATA_VERIFY_STEPPED=1`: - **Segmented capture.** The window is captured as graph segments, each ending at a layer's doorbell publish. - **CPU experts between segments.** The host serves the CPU experts between segments, then raises the flag before it submits the segment that consumes them. - **Pipelined helper.** Plan copy and VRAM groups run on a helper thread pinned to the host thread's SMT sibling. `STRATA_VERIFY_PIPE=0` and `STRATA_PIPE_HELPER_SIBLING=0` turn the pipelining and the pinning off. - **Cached host buffers and plain payload stores.** In this mode no kernel reads a host store while it runs, and the host reads a segment's rings and payload only after that segment has drained. So the window's mapped buffers come from ordinary `malloc_host` instead of `host_malloc_polled`, and the doorbell publishes use plain payload stores instead of `sys_store_mapped`. The payload checksum still guards every host read. - `STRATA_VERIFY_HOST_UNCACHED=1` keeps both uncached. - With `STRATA_VERIFY_STEPPED` unset, every buffer and store is as on `main`. ## Measured Setup: - Hardware: Arc A770 16 GB (xe driver, oneAPI 2025.3, Linux 7.0), Ryzen 9 3900X, 32 GB DDR4. - Model: Qwen3-Coder IQ1_M, learned24 expert profile, `--resident-experts --spec 4 --mtp`, greedy decoding, 256 new tokens. - Prompt: code, 7,987 tokens. - The runs are not interleaved. RAM headroom differs where noted, because resident experts are mlocked and the box has 32 GB. | Build | Decode tok/s | GPU wait for the CPU per window | |---|---|---| | #1602 head (`712bbc9`) | 12.3 | 170 ms | | this PR, `STRATA_VERIFY_STEPPED=0` | 12.6–13.1 | 134 ms | | this PR, stepped, everything uncached (= `STRATA_VERIFY_HOST_UNCACHED=1`, measured on the dev branch) | 21.3–25.0 | 70–72 ms | | this PR, stepped (dev branch, same commit) | **31.8–32.5** (4 runs) | 42.7–43.0 ms | | this PR, stepped, 32K-token prompt (dev branch) | 27.2 | 41.9 ms | The dev branch is this commit on top of an A770 rebase. Its resident set was 17.9 GiB, against 13.7 GiB for this branch with `STRATA_RESIDENT_HEADROOM_GIB=6`. On this PR branch alone, with the smaller resident set, stepped decode was 19.9 tok/s. It was memory-bound because part of the expert set was not resident. Output check (greedy): this PR stepped, this PR stepped-off, and #1602 agree for the first 80 tokens. Run-to-run, #1602 itself agrees for 76–79 tokens on this prompt, so that is the existing nondeterminism, not a change. ## Not tested - Arc A750/i915, B-series, and multi-GPU. - The CUDA and HIP builds. The only shared-header change is one new SYCL-only declaration, `doorbell_plain_payload`, in `include/strata/kernels/elementwise.hpp`; its definition is in `sycl/`. On exit, #1602 and this branch both sometimes segfault after the stats are printed (rc 139). It also happens without this commit.
Sur le site
Liens install, modèles, releases.